Skip to content
Swizztool

OCR PDF - Scan to Text

Turn scanned PDFs and photos into editable, searchable text. The language file downloads once. Text never leaves the device.

Drop a scanned PDF or image

Text is read right in this browser. The English language file downloads once. Handwriting and low-contrast photos read worse than clean printed scans.

How to use OCR PDF

  1. Drop a scanned PDF or a photo of a page.
  2. Pick a language. The first run downloads that language file.
  3. Run OCR and wait while each page is read locally.
  4. Edit, copy, or download the recognized text.

About OCR PDF - Scan to Text

Turn scanned PDFs and photos of documents into editable text. Choose English, Spanish, French, or German, run text recognition right in your browser, and edit, copy, or download the result.

Related tools: Extract Text from PDF, PDF to JPG Converter, Word Counter, Free Readability Checker

OCR or text extraction

Extract PDF Text copies text that is already stored in a digital PDF. OCR (optical character recognition) reads the shapes of letters in an image, which is what you need for scans, photos, and faxes. If you can highlight text in your PDF reader, use Extract PDF Text instead; it is faster and exact.

The first time you use a language, its recognition data (a few megabytes) downloads once and is then reused, so later runs start straight away.

Getting the best results

Recognition is most accurate on clear, straight, printed text. Photograph pages flat and well lit, crop away the background, and use the highest resolution available. Handwriting, unusual fonts, and complex tables are recognized less reliably.

Always proofread the result, especially numbers, names, and anything you will rely on.

Common uses

Scanned letters
Get editable text from a scanned letter or contract.
Photos of documents
Copy text from a photo of a printed page.
Old records
Digitize printed documents for search and editing.
Quotes from books
Copy a passage from a photographed page.

Questions about OCR PDF

Does OCR upload my scan?
No. Tesseract.js and PDF.js run in this tab. The language file downloads from a CDN once and stays cached.
When should I use extract text instead?
If you can already highlight the text in a PDF reader, use Extract PDF Text. It is faster and more accurate than OCR.
How accurate is the text recognition?
It is strong on clean printed English and other supported languages. It is weaker on handwriting, photos taken at an angle, and dense tables.
Why is the first run slow?
The worker and language data are a few megabytes. After that, later pages and later visits reuse the cached model.
Does OCR work on handwriting?
Poorly. It is designed for printed text.
Is my scan uploaded?
No. Recognition runs in your browser. Only the language data is downloaded.

Comments

Questions, ideas, or a bug in OCR PDF? Let us know.