The difference between page splitting and category splitting
A normal PDF splitter asks which page numbers you want. Category splitting instead looks for text that repeats across pages and may represent a meaningful label such as a branch code.
The difficult part is deciding which repeated text is actually a category. Page headers, dates and report titles also repeat, so automatic detection must present candidates rather than silently assuming every recurring phrase is useful.
How FileFiddler approaches it
Page text is extracted locally. Lightweight page text is sent to the protected FileFiddler service for category analysis, where recurring short labels can be suggested. You choose which labels represent your real categories before the final split is produced.
The actual PDF pages remain in the browser and are split locally. This separates the classification step from the document-building step.
Sources that work best
The tool is less reliable when category names change spelling from page to page, when the only labels are inside scanned images, or when the same page legitimately belongs to multiple categories.
- Reports with a consistent branch or department code on each page.
- Statements where an account type appears in a stable header position.
- Batch reports produced from one system using a repeated template.
Review before download
Always review the proposed category labels before producing final files. A recurring report title is not a branch, and an accidental match can put pages into the wrong output group.
For a one-time simple extraction, the normal Split PDF tool may be faster. Category splitting becomes valuable when the document is large and its labeling pattern is consistent.
Related FileFiddler tools
About this guide
This guide documents how FileFiddler approaches the underlying file problem, including limitations that matter when you review the output. It is maintained alongside the tools rather than written as generic promotional content.
