OCR PDF: make scans searchable
Runs on your device- Searchable, copyable output
- Original pages stay untouched
- Runs on your device
- Free, no sign-up
or drop it here · up to 100 MB
Processed on your device. Nothing is uploaded.
Recognize the text in scanned pages and add an invisible text layer you can search, select and copy, while the pages look exactly as before.
How it works
How to make a scanned PDF searchable with OCR
- 1
Add a scanned PDF
Drop a PDF made from a scanner, phone scan app or fax onto the page. Files up to 100 MB are accepted. Enter the password first if the file is protected.
- 2
Choose which pages to read
By default, pages that already contain text are skipped. Tick the option to OCR them anyway if their existing text is wrong or incomplete.
- 3
Make it searchable
Press Make searchable. Each page is rendered at about 250 DPI and read on your device, with progress shown page by page and a Cancel button.
- 4
Download and search
Open the new PDF and use Find (Ctrl+F or Cmd+F): the recognized words sit invisibly on top of the original scan.
Optimize
About OCR PDF
Last updated
When does a PDF need OCR?
A PDF needs OCR when you can't select or search its text, which means the pages are only pictures of paper. You can read such a scan, but your computer can't: Find returns nothing, you can't copy a paragraph into an email, screen readers stay silent, and document management systems can't index it. OCR fixes that for contracts, invoices, old letters, book chapters, receipts and anything that came out of a scanner or phone scan app.
If you can already select text in a PDF, it doesn't need OCR. That's why any page that already holds 50 or more text characters is skipped automatically, which keeps mixed documents fast and avoids a duplicate text layer. Tick the option to OCR those pages anyway when their existing text is garbled, for example after a poor earlier OCR run.
How does OCR PDF work?
OCR PDF renders each page as an image at about 250 DPI with pdf.js and passes it to Tesseract, an open-source OCR engine running as WebAssembly in your browser. Tesseract's neural-network recognizer returns every word with its position, baseline and a confidence score, and words scoring below 15 out of 100 are discarded as noise.
Those positions are mapped back onto the original PDF page, taking page rotation and cropping into account, and the words are written as invisible text exactly over the printed ones. Because the original page content is kept rather than replaced with a new image, the scan's quality and file structure stay as they were, and the file grows only by the size of the text, typically a few kilobytes per page.
Should I make a searchable PDF or convert to Word?
Make a searchable PDF when the document must keep its exact appearance, and convert to Word only when you need to rewrite the content. A searchable PDF keeps the scan exactly as it looks, which matters for signed contracts, forms, certificates and anything you may need to prove later. It is also the format most archives, courts and document systems expect.
In a Word conversion the layout is rebuilt from the recognized text, so tables, columns and fonts rarely match the original perfectly. A common workflow is to OCR the scan here, keep that PDF as the record, and send a copy through PDF to Word for editing. When only the raw wording is needed, PDF to Text extracts the new text layer into a plain file.
How do I get better recognition?
Better recognition comes mostly from a better scan, because the engine can only read what is clearly visible on the page. Sharp, straight, evenly lit pages with dark text on a light background are read almost without errors, while shadows, curved book pages and heavy JPEG artefacts cause misread letters. These steps help the most:
- Scan at 300 DPI in greyscale or black and white. Lower resolutions lose small print.
- Keep pages straight and evenly lit. Phone scan apps that flatten and crop the page work much better than plain photos.
- Rotate sideways pages with Rotate PDF before running OCR.
- Large photo scans can be shrunk with Compress PDF first; OCR works on the page as it looks, so a moderate recompression rarely hurts accuracy.
What are the limits of OCR PDF?
OCR PDF reads printed English only, and it does not recognize handwriting reliably. Complex layouts such as multi-column tables may copy out in an unexpected order even when every word is found. The invisible layer uses the standard Helvetica font, so ligatures and unusual dashes are simplified to plain letters and hyphens, and characters outside the Western character set are stored as question marks.
Files up to 100 MB are accepted, and very long scans are best split into parts on phones, where memory is tighter. If no words are found at all, the scan is probably blank, too faint, below about 150 DPI, or not in English. Recognition can be cancelled at any point, and nothing is saved until the final page has been read.
FAQ
OCR PDF questions
What does OCR do to a PDF?
It turns pictures of letters into real text. Here that text is added as an invisible layer aligned with the words on the page, so the scan looks unchanged but you can search, highlight and copy it.
Do scanned documents leave my device during OCR?
No. The recognition engine (Tesseract) and the English language model, about 3 MB, are downloaded from this site and run inside your browser. The PDF itself stays on your device from start to finish.
Which languages are supported?
English only at the moment. Other Latin-script languages are partly recognized, but accents and special letters may come out wrong, and scripts such as Chinese, Arabic or Cyrillic are not supported yet.
How long does OCR take?
Roughly 5 to 20 seconds per page, depending on your device and how dense the page is. A 20-page scan usually finishes in two to seven minutes on a laptop, and phones are slower.
How accurate is the text recognition?
Very accurate on clean, straight scans at 300 DPI. Handwriting, faint photocopies, skewed phone photos, tables and decorative fonts produce more errors, so check important numbers such as totals and account details before relying on them.
Can I edit the text after OCR?
Not in place. The OCR layer makes text searchable and copyable, not editable. To edit the wording, run PDF to Word on the searchable PDF, or use Edit PDF to add new text on top.
Does OCR work on a password-protected PDF?
Yes, once you enter the password. The pages are read as usual, but the searchable copy is saved without the password. Add protection again afterwards with Protect PDF if the document still needs it.
Keep going
Related tools
Got a PDF to fix?
Free, no account, no watermark. Most tools run on your device, so files never leave it.
Browse all tools