SandboxPDF

PDF to text

Just the words. Ready for anything.

Extract selectable text as a UTF-8 file. Image-only pages need OCR first.

How to use PDF to text

  1. Open your document from your device.
  2. Adjust the settings and preview the result.
  3. Read text results on screen or download your file. Your original remains untouched.

Local by design

Your documents are processed on your device. No account, payment or document upload is required.

Enable JavaScript to use the document workspace.

Capabilities and limitations

Extract the words for reading, searching and reuse

PDF to text saves selectable document text as a UTF-8 file. It is useful for notes, searching with other tools or preparing text for a separate editing workflow. It does not preserve the page’s visual design, and it cannot automatically extract words from an image-only scan without recognition.

Inspect the source first. Try selecting a word in a PDF reader and look at a page with columns, tables or footnotes. A PDF can display correctly while storing text in an order that differs from normal reading. The output should therefore be checked as prose rather than assumed correct because every visible page looked polished.

Create a text copy you can inspect

  1. Open the PDF and choose the relevant page range.
  2. Confirm that the selected pages contain a usable text layer.
  3. Choose an output name that remains connected to the source.
  4. Create and download the UTF-8 text file.
  5. Open it in a text editor and compare representative passages with the PDF.

Page selection uses physical positions in the PDF. Front matter can make those positions differ from printed page numbers. If the text will support a citation or a review, record the selected range clearly. A plain-text copy may not preserve every navigational cue that made the original page easy to reference.

Check reading order and character mapping

Columns can become interleaved, headers can interrupt paragraphs and table values can lose their labels. Read a complete paragraph from a difficult page and verify that it starts and ends correctly. Inspect a table separately rather than treating a sequence of numbers as a reliable dataset without column relationships.

Some documents omit literal spaces and position characters geometrically. Extraction may need to infer word boundaries, producing joined or broken words. Unusual font encodings can also yield incorrect characters even when the page renders properly. Check accents, ligatures, mathematical symbols and currency signs in the saved text.

Line-end hyphenation needs context. Joining every broken line can damage identifiers or genuine compound words, while preserving every break can make prose awkward. If you edit the extracted text, distinguish your cleanup from exact quotation. Keep the source available so important wording can be verified.

Use OCR when the source is a scan

An image-only page contains pixels, not selectable words. Run an appropriate OCR operation first if text extraction is required. Then check the recognised output against the image, especially names, dates and amounts. OCR adds another interpretation stage and is not guaranteed to reproduce every character correctly.

A text file is useful for many downstream tasks, but it does not preserve accessibility structure, interactive forms, images or document authenticity evidence. It is not a substitute for an original signed record. If a summary or translation will use the text, inspect extraction quality before invoking the next tool; later language processing cannot reliably repair missing context.

For a large document, begin with a representative range and confirm that the format is useful. Use the reading-order guide for difficult layouts and OCR quality checks for scans. Save the original PDF and text derivative separately, and verify the exact downloaded text rather than relying on a page preview.

Frequently asked questions

Why is the text file empty for a visible scan?

The page may contain only an image. Recognition is needed to create machine-readable text.

Is the original formatting preserved?

No. This output focuses on words in plain text, not the PDF’s page design.

Why do columns copy in the wrong order?

PDF drawing order can differ from reading order. Complex layouts need review and sometimes manual correction.

Does UTF-8 support non-English characters?

Yes, but correct output also depends on the PDF’s character mapping or OCR quality. Inspect the saved characters.

Can I treat the output as an exact transcript?

Only after checking it against the source. Extraction errors and layout interpretation can change the apparent wording or relationships.