The quickest test
Open the PDF and try to select a single word with your mouse or trackpad. If individual words can be highlighted and copied, the document is probably text-based. If dragging only selects an entire page image, or nothing useful can be copied, it is probably scanned.
Some files are mixed: a cover page may be scanned while later pages contain selectable text. A converter may therefore behave differently from page to page.
Why text-based PDFs are easier
A text PDF exposes characters and positions directly. FileFiddler can use those positions for table reconstruction or use detected text runs when building an editable Word document. The original fonts and layout still may not translate perfectly, but the source information is much richer.
Scanned documents require optical character recognition. OCR estimates which pixels represent letters and numbers, then creates text boxes with confidence and coordinates. That extra recognition step introduces possible errors before table reconstruction even starts.
What affects OCR quality
Clear, upright images with readable type produce much better results than blurred photographs, skewed pages or compressed screenshots. Small decimal points, commas and similar-looking characters can be especially difficult.
If a scan is faint or rotated, improving the image first can be more effective than repeatedly converting the same source.
- Use the highest-quality source you have.
- Crop large empty borders when practical.
- Correct obvious rotation before OCR.
- Review totals, dates and account numbers after recognition.
Choosing the FileFiddler workflow
Use PDF to Excel when the PDF contains selectable table text. Use Image to Excel when the table exists in a photo, screenshot or scanned page. Use PDF to Word when you need editable document text rather than a spreadsheet.
The important question is not simply “Is this a PDF?” but “What information is actually stored inside this PDF?”
Related FileFiddler tools
About this guide
This guide documents how FileFiddler approaches the underlying file problem, including limitations that matter when you review the output. It is maintained alongside the tools rather than written as generic promotional content.
