Free · Fast · Privacy-first

Scanned PDF to Excel

A scanned PDF is fundamentally an image: each page is a flat picture of what was originally a printed sheet, with no machine-readable text underneath.

OCR + table extraction

🔒

Works on common scan layouts

Browser-based privacy

Manual review supported

Cost
Free tier
Sign-up
Not required
Processing
Tool-specific
Privacy
Clearly disclosed
IframeResponsiveAttribution included

Add this PDF to Excel to your website

Drop the PDF to Excel into a blog post, product docs, intranet, or school portal with one iframe. Processing, privacy, and usage limits are the same as on the full tool page.

  • One copy-ready line of HTML
  • Responsive — adapts to any container width
  • No API credentials are placed in the snippet

Embed code

<iframe
  src="https://www.fixtools.io/pdf/pdf-to-excel?embed=1"
  width="100%"
  height="780"
  frameborder="0"
  style="border:0;border-radius:16px;max-width:900px;"
  title="PDF to Excel by FixTools"
  loading="lazy"
  allow="clipboard-write"
></iframe>

Attribution-friendly: a small "Powered by FixTools" link appears in the embed footer.

OCR quality is the dominant factor in scanned PDF conversion accuracy

Optical character recognition works by analysing the shapes of glyphs in an image and matching them to characters in a learned model. Modern OCR engines like Tesseract and the engines built into Adobe Acrobat achieve accuracy above 99 percent on clean 300 DPI scans of standard fonts. The same engines drop below 90 percent on 150 DPI scans, low-contrast photocopies, or skewed phone photos. For a table with 500 numeric cells, the difference between 99 percent and 90 percent accuracy is the difference between 5 errors and 50 errors. Those errors are not evenly distributed: digits zero and capital O are commonly confused, as are digit one and lowercase L, which means numeric SKU columns suffer disproportionately on poor-quality scans.

FixTools detects when an uploaded PDF lacks a text layer and either runs OCR in the browser or directs you to the OCR PDF tool to produce a searchable version first. Browser-based OCR is feasible for short documents but slow for long ones because the inference happens on your device CPU. For documents over 20 pages, running OCR through a desktop tool like Adobe Acrobat Pro or a server-side OCR service produces faster results, after which the searchable PDF can be converted to Excel using the standard FixTools converter. The two-tool workflow trades a small amount of friction for substantial speed gains on longer documents.

Scan quality factors that materially affect conversion accuracy include resolution (300 DPI minimum, 600 DPI ideal for small fonts), skew (pages should be straight, not tilted), contrast (sharp black text on white background), and brightness (consistent lighting across the page). Phone photos taken in office lighting are often acceptable for casual reading but produce notably worse OCR results than flatbed scans because of uneven lighting, slight perspective distortion, and lower effective resolution. If you have access to a flatbed scanner or a multifunction printer with a document feeder, scanning at 300 DPI in black-and-white mode produces dramatically better OCR results than phone capture.

After OCR-based conversion, manual review of the output is essential for any data where accuracy matters. Sort the converted columns to find values that look anomalous: numbers with too many digits, dates in the wrong format, or text strings where you expect numerics. These anomalies usually flag the rows where OCR misread something. A short review pass that catches half a dozen errors on a five-hundred-row extraction is far faster than retyping the entire table from scratch and produces clean data with high confidence in its accuracy.

How to use this tool

💡

Upload your scanned PDF. The tool will OCR the document if no text layer exists, then extract tables into Excel. Review the output for OCR errors.

How It Works

Step-by-step guide to scanned pdf to excel:

  1. 1

    Check scan quality first

    Before uploading, open your scanned PDF and check the image quality. Pages should be straight, text should be clearly legible at 100 percent zoom, and contrast should be sharp. If quality is poor, re-scan at 300 DPI minimum if possible. Better source scans produce better OCR which produces better Excel output, so investing a minute in re-scanning saves much more time downstream.

  2. 2

    Upload to FixTools

    Drag your scanned PDF onto the upload area. The tool detects that the file lacks a text layer and prompts you to run OCR first. For short documents the OCR can run in the browser; for longer documents the tool recommends using the dedicated OCR PDF tool or an external OCR service before converting.

  3. 3

    Run OCR if needed

    If prompted, run OCR on the document. This adds a searchable text layer to the PDF without changing its visual appearance. The OCR step typically takes ten to thirty seconds per page on a modern device. Once OCR completes, the PDF is ready for table extraction using the standard converter.

  4. 4

    Convert and review

    Convert the now-searchable PDF to Excel using the standard converter. Open the output XLSX and review for OCR errors: sort columns to find anomalous values, scan for obvious misreads in alphanumeric fields, and verify totals against the source document. Manual correction of a handful of errors is much faster than retyping the entire table.

Real-world examples

Common situations where this approach makes a real difference:

Lawyer digitising historical case records

A solicitor inherits a case file with twenty years of scanned correspondence and ledger pages that need to be loaded into a modern case management system. Converting the scanned ledger pages to Excel through OCR makes the historical data searchable and sortable for the first time, enabling timeline reconstruction that would have been impossible to do manually within the budget of the case.

Operations manager digitising paper inventory logs

A warehouse operations manager finds boxes of handwritten and typed inventory logs from before the company adopted electronic record-keeping. Scanning each page at 300 DPI and converting through OCR to Excel allows the historical inventory data to be loaded into the current ERP for trend analysis and audit support, recovering institutional knowledge that would otherwise have stayed locked in paper.

Researcher extracting historical census data

A demographer researching long-term population trends needs to extract tables from scanned historical census reports going back fifty years. The scanned PDFs have been digitised by archive services but never converted to structured data. Running each scanned PDF through OCR and Excel conversion produces machine-readable tables that can be loaded into statistical software for the kind of longitudinal analysis the source data was designed to enable.

Accountant working with paper supplier invoices

An accountant for a small business receives some supplier invoices on paper that have been scanned and emailed as PDFs. Converting these scanned invoices through OCR to Excel produces line-item data that can be coded and posted alongside the electronically-delivered invoices, eliminating the need to retype paper-sourced bills as a separate manual workflow.

Pro tips

Get better results with these expert suggestions:

1

Re-scan at higher resolution if accuracy is poor

If your first OCR-and-convert pass produces a noisy output with many recognition errors, the cheapest improvement is to re-scan the source document at higher resolution. Going from 150 DPI to 300 DPI typically cuts OCR errors by a factor of three or more. Going from 300 to 600 DPI is incremental but worth the time for small-font financial documents. Better source data is almost always faster to obtain than fixing errors downstream.

2

Use black-and-white scan mode for text-heavy pages

Scanning in true black-and-white (1-bit) mode produces sharper character edges than greyscale or colour scans for documents that contain only text and tables. Sharp edges give OCR engines cleaner shape signals, which directly improves recognition accuracy. Reserve greyscale or colour scanning for pages where photos or colour content matter to the document's meaning; for pure text and tables, black-and-white is strictly better for OCR.

3

Filter for common OCR errors after conversion

Search the converted Excel for characters that OCR commonly confuses: lowercase l and digit 1, capital O and digit 0, capital S and digit 5, lowercase i and lowercase l. In numeric columns, these confusions often produce values that are obviously wrong (an extra digit, an unexpected letter). A find-and-replace pass on each suspicious pair cleans up the majority of OCR errors in a few minutes.

4

Crop pages tight before scanning

Black borders, edge shadows, and stray marks at the page margins distract OCR engines and can cause spurious character recognition. Crop your scanned pages tight to the actual document content before running OCR. Most scanning software supports automatic edge detection that achieves this in one click. A cleanly cropped page produces cleaner OCR which produces cleaner Excel output.

FAQ

Frequently asked questions

Yes, with OCR. Scanned PDFs are images of pages with no machine-readable text, so the converter requires a text layer to extract structured data. FixTools detects when a file lacks a text layer and prompts you to run OCR first. After OCR adds searchable text, the standard table extraction works the same way as for digitally-generated PDFs. Quality depends heavily on the source scan: 300 DPI clean scans produce near-perfect results, while low-resolution phone photos may need manual cleanup.
Modern OCR engines achieve 99 percent or better accuracy on clean 300 DPI scans of standard fonts, which is the resolution typical banks and accounting systems produce when they print and scan documents. Accuracy drops with scan resolution and with non-standard fonts. Always verify extracted totals against the source document; matching totals confirm the body cells extracted correctly, while mismatches indicate cells that need manual correction.
Yes. Adobe Acrobat Pro, ABBYY FineReader, and various open-source tools can produce a searchable PDF that FixTools then converts to Excel without needing to run OCR itself. This separation is useful for large documents where browser-based OCR would be slow. Run OCR in your preferred tool, save the resulting searchable PDF, and feed that file into the FixTools converter for the table extraction step.
A minimum of 300 DPI in black-and-white or greyscale, with pages straight (not tilted) and contrast sharp. 600 DPI is ideal for documents with small fonts. Phone photos can work in a pinch but typically produce noticeably worse OCR than flatbed scans because of uneven lighting and perspective distortion. If you have access to a multifunction printer with a document feeder, that produces ideal scans for OCR.
For browser-based OCR, expect roughly ten to thirty seconds per page on a modern device. A twenty-page document takes a few minutes. For longer documents, using a desktop tool like Adobe Acrobat Pro or a server-side OCR service produces results faster because those tools can parallelise across CPU cores more aggressively than a single browser tab. The total end-to-end time depends on both OCR speed and the table extraction step that follows.
Standard OCR engines are tuned for printed text and perform poorly on handwriting. For handwritten cells you typically need to fall back to manual transcription or use specialised handwriting recognition (HWR) tools. If your scanned document mixes printed and handwritten content, OCR will extract the printed portions accurately and the handwritten cells will appear as blank or garbled in the output, which you can fill in by hand.
Yes. OCR and table extraction both run in your browser tab without transmitting file content to any server. Your scanned documents, including any sensitive information they contain, stay on your device throughout the process. This is particularly important for medical records, legal documents, and financial statements where any third-party exposure could create compliance issues.
If OCR fails to recognise text in a particular column (which can happen with very small fonts or low-contrast print), the column will appear empty in the output Excel. The fix is usually to re-scan the source at higher resolution and re-run OCR. If re-scanning is not possible, you can manually type the values for that column referring back to the source PDF, which is still much faster than retyping the entire table.

Related guides

More use-case guides for the same tool:

Ready to get started?

Open PDF to Excel to review its free limits and processing method.

Open PDF to Excel →

Free tier · No account needed · Transparent limits