What a PDF actually stores
A spreadsheet normally stores a value in a cell at a known row and column. A PDF can instead store many small text objects positioned at coordinates on a page. Two values that appear to sit in the same visual column may have no explicit relationship in the PDF file itself.
That is why copying a PDF table can produce unexpected line breaks, merged text, missing spaces or values in the wrong order. The page looks structured to a person because our eyes infer rows, headings and alignment. Software has to infer those relationships from the document.
What FileFiddler reconstructs
For text-based PDFs, FileFiddler reads selectable text together with its position on the page. The table reconstruction step groups content into likely rows, estimates column boundaries and builds spreadsheet cells from that geometry.
This works best when the source uses consistent alignment. Supplier statements, price lists, ledgers and many business reports are good examples. Irregular multi-column layouts, wrapped descriptions and decorative documents are harder because the visual page does not map cleanly to a rectangular worksheet.
- Selectable text is preferable to a scanned page.
- Consistent left edges and numeric columns improve reconstruction.
- Repeated headers can help identify page structure but may need cleanup in the workbook.
- A PDF can look like a table while containing no real table object at all.
Why manual column guides matter
Automatic detection cannot know every document layout. A total column may sit very close to a quantity column, or a long description may visually overlap the next field. FileFiddler therefore lets you review and adjust column boundaries instead of forcing one automatic answer.
Moving a column guide tells the reconstruction step where one field should end and the next should begin. This is particularly useful for recurring reports where the source layout is consistent but unusual.
When to use OCR instead
If you cannot select the text in the PDF, the page may be an image or scan. In that case there are no useful PDF text coordinates to reconstruct. Render or export the page as an image and use the Image to Excel workflow, which first performs OCR and then reconstructs the table from recognized text boxes.
Choosing the correct workflow matters more than choosing a higher conversion setting. Text extraction and OCR solve different source problems.
Related FileFiddler tools
About this guide
This guide documents how FileFiddler approaches the underlying file problem, including limitations that matter when you review the output. It is maintained alongside the tools rather than written as generic promotional content.
