OCR PDF
Make scanned words searchable.
Recognize scanned text in English, French, German or Spanish. Downloads a language model on first use; recognition runs locally. Proofread the result.
How to use OCR PDF
- Open your document from your device.
- Adjust the settings and preview the result.
- Read text results on screen or download your file. Your original remains untouched.
Local by design
Your documents are processed on your device. No account, payment or document upload is required.
Enable JavaScript to use the document workspace.
Capabilities and limitationsMake scanned words searchable with a reviewable text layer
OCR PDF recognises words from rendered pages using local Tesseract models. It is useful for searching scanned correspondence, copying a passage from a document photograph or preparing text for later reading tools. Recognition is an interpretation of pixels, not a recovery of the original editable source.
SandboxPDF currently offers OCR in English, French, German and Spanish. That language list is separate from summary and translation support. Choose the appropriate recognition language and inspect mixed-language material carefully. Names, codes and uncommon vocabulary may not follow the model’s usual word patterns.
Start with a representative page
- Open a readable scan and correct gross orientation if necessary.
- Select a manageable page range and the appropriate OCR language.
- Choose the available rendering resolution, balancing detail and device memory.
- Allow the first-use model download, then create the result.
- Reopen the output, search a known phrase and compare copied text with the image.
The model download supplies recognition assets; it does not upload the PDF for remote recognition. Keep the tab open during processing. Long, high-resolution scans can consume substantial memory, especially on a phone. Test a small range before committing to a large document and avoid competing heavy tasks.
Improve the source before trusting the output
Even lighting, sharp focus and upright text help recognition. Shadows, curved book pages, strong perspective distortion and faint printing can reduce accuracy. Increasing pixel dimensions does not recreate detail lost in a blurry photograph. If the paper is available, a better capture can be more effective than repeated recognition attempts.
Inspect the weakest pages rather than only the clear cover. Small footnotes and dense tables can fail differently from large paragraphs. A language model may turn an unusual word into a plausible but wrong one, so smooth-looking output is not enough. Check names and technical terms directly against the image.
Numbers deserve a separate review. Zero and O, one and l, decimal separators and minus signs can be confused. Verify amounts, dates and identifiers character by character when they matter. For tables, check that the correct value remains associated with the correct label; accurate characters in the wrong order are still incorrect information.
Plan downstream operations carefully
Searchable text does not automatically make a document fully accessible. Headings, table relationships and reading order may need additional authoring or remediation. Likewise, OCR does not recover spreadsheet formulas or original Word styles. It provides recognised text that can support other workflows after verification.
If redaction is required, recognise only the verified redacted page images. Reusing a source text layer can reintroduce information that no longer appears visually. If you later use image-mode compression, check whether the recognised text layer has been removed by that visual rebuild.
Before summarising, translating or chatting with the document, inspect extraction quality. Those tools cannot reliably correct every recognition error merely because their responses sound natural. Read OCR quality verification and PDF reading order for a fuller review. Keep the scan as the source of truth and verify the final downloaded file after the complete sequence of operations.
Frequently asked questions
Does OCR upload my scan?
No. Recognition runs locally after the required language assets download.
Why are some words wrong?
Recognition depends on image quality, language and layout. It can misread even a visually plausible page.
Are Arabic and Chinese OCR offered here?
The current OCR controls list English, French, German and Spanish. Other tools have separate language capabilities.
Does OCR make a PDF fully accessible?
No. Searchable text is only one part of accessibility; meaningful structure and reading order need separate review.
What should I verify first?
Search a known phrase, copy a paragraph and check critical numbers and names against the source image.