SandboxPDF

Why scanned PDFs are so large—and how to make a smaller copy safely

A scan can look like a simple page of text while storing millions of coloured pixels. Before choosing a compression setting, identify what the file contains and decide which details the recipient actually needs.

A page of words can really be a photograph

A document exported from a word processor can describe a sentence using characters, font information and positions. A scanner usually records a rectangular image instead. The PDF is then a container around that image. Twenty pages with only a few paragraphs can therefore be larger than a long report created directly from editable text. The apparent amount of writing is a poor guide to the number of bytes involved.

Try selecting an individual word in your PDF reader. If selection behaves like a single photograph, the page may have no text layer. If words are selectable, it might be a native digital document, or a scan with an invisible OCR layer. Zooming in helps distinguish the two: a scanned letter eventually reveals pixels, while a vector letter stays smooth. Some files combine all these types on different pages, so inspect more than the cover.

This distinction matters because a compressor cannot create the original editable document from a photograph. OCR can recognise words, but the page image still occupies space. A smaller file produced by discarding that image would be a different representation of the document, with its own risks. Keep the original scan whenever stamps, handwriting, signatures, photographs or layout carry meaning.

Resolution grows in two directions

Doubling a scan’s resolution doubles its horizontal and vertical pixel counts. The total number of pixels increases roughly fourfold. A full-page image at 600 dots per inch consequently contains far more information than the same page at 300 dots per inch, even when you cannot see a difference at ordinary reading size. The final file size also depends on the image encoding, noise and colour content; it will not follow one exact ratio.

Higher resolution can be worthwhile for fine print, line drawings or later recognition. It is less useful when the source itself is blurred or photographed out of focus. Enlarging a blurry scan does not recreate missing detail. If you have access to the paper, a sharp replacement capture may produce both better readability and a more compressible image than aggressive processing of the damaged scan.

Do not treat a single resolution as correct for every purpose. A screen-reading copy, a small-print contract and a technical drawing have different needs. The useful test is whether the smallest important detail survives at the intended reading or printing size. Check punctuation, decimal separators, superscripts, shaded table headings and faint handwritten additions—not just the largest heading on page one.

Colour, paper texture and camera noise also matter

A page that looks white may contain subtle shadows, yellow paper, compression artefacts and sensor noise. Those variations make an image more complex. Phone photographs can include a desk, fingers, deep gutter shadows and an uneven background. Cropping and improving capture conditions can remove irrelevant visual complexity before you start adjusting quality settings.

Grayscale is often appropriate for plain correspondence, but it can erase useful distinctions. A red correction, a coloured legend or two chart lines may become difficult to distinguish. Black-and-white thresholding is more aggressive: faint writing can disappear entirely. Always inspect a representative page with the most delicate detail. If the document mixes photographs with text, one setting may not suit every page equally well.

Avoid assuming that a background-cleaning filter improves evidence. A faint stamp or pencil note may be precisely the information someone needs. For an administrative reading copy, reducing shadows may be helpful; for a record that must faithfully preserve the original appearance, keep the unmodified capture and label any processed version clearly.

Start with a lossless attempt

Open Compress PDF and try lossless compression first. In SandboxPDF, this rewrites and recompresses supported PDF streams rather than turning every page into a new image. It is the sensible starting point when you want to preserve selectable text and avoid intentional visual degradation. Download the result and compare both its size and its behaviour with the source.

A lossless result may barely shrink. That is not necessarily a failure. The images might already use an efficient encoding, or the document may contain little redundant structure. A small increase is also possible when the rewritten container has different overhead. Judge the exported bytes rather than assuming a successful completion message means the file became smaller.

If the file already meets the recipient’s limit, stop. Repeated compression adds work and can make version management confusing. Save the original and one clearly named delivery copy. If it remains too large, decide whether to reduce image quality, remove genuinely unnecessary pages, or use a recipient-approved transfer method. These choices solve different problems and should not be hidden behind a single percentage claim.

Understand the trade-off in image mode

SandboxPDF’s image compression mode renders pages and creates a flattened image-based PDF. This can substantially reduce some scans, but it also removes selectable text, live links and interactive form fields from the resulting copy. Even a native digital document becomes a set of page images. The visible page may look familiar while important document behaviour has changed.

Use this mode when a visual reading copy is sufficient. Avoid it when the recipient needs to search, copy text, complete a form, inspect a digital signature or use accessibility structure. If search is important, you may need OCR afterwards, followed by proofreading. That adds a new recognised text layer; it does not recover the original document structure or restore every feature discarded during flattening.

Choose moderate settings first and inspect the result at normal size and enlarged. If a table’s numbers become ambiguous, raise the quality or resolution. If photographs look acceptable but small labels do not, preserve the labels rather than optimising only for the photographs. A smaller file that changes a value is a failed delivery copy, regardless of how impressive the size reduction looks.

Work through a practical example

Imagine a twenty-page application packet that must fit below a portal’s upload limit. It includes eighteen scanned text pages and two pages of colour diagrams. First, record the original size and page count. Confirm whether the portal specifies a maximum per file or for the whole submission, and whether it requires one combined document. Those details affect whether splitting is an acceptable option.

Next, try lossless compression and inspect the exported size. If it still exceeds the limit, create a separate image-mode copy. Examine the diagram legends, the smallest text and any signatures. Check that all twenty pages remain in order. Search the result for a known phrase if searchability is required; a visual-only export will not pass that check without a text layer.

If the diagrams become unreadable before the file reaches the limit, do not keep lowering quality. Consider extracting the diagrams into a separate file if the portal permits it, or returning to the original scan settings. You might also discover duplicate cover sheets or blank backs that can be removed without affecting the submission. Verify the recipient’s rules before changing the packet’s structure.

A useful comparison takes more than opening the file

Review the original and exported versions side by side. Count pages, compare orientation and look at the edges for clipping. Read several identifiers character by character. Inspect a faint page and a dense page as well as an easy one. If the file includes a barcode or QR code, test it with the intended reader; visual similarity alone does not prove that a machine-readable symbol still works.

For printing, produce a small proof at the intended scale. A scan can look satisfactory when fitted to a bright display but become muddy on paper. Printer scaling can add another reduction, especially when A4 and Letter are mixed. If you are preparing evidence, an archive submission or a formal filing, follow that organisation’s requirements rather than this general reading-copy workflow.

Keep the delivery filename distinct, such as an original basename with a “compressed” suffix. Reopen the actual downloaded file, not only the in-browser preview. This catches accidental selection of the wrong result and confirms that the saved copy is usable. Preserve the original separately so that a later request for better quality does not force you to work from the compressed derivative.

When compression is the wrong next step

If the document contains unreadable scans, improve capture or recognition first. If it contains confidential information, handle redaction before distribution. If it contains a signed certification, ask whether rewriting the file is permitted; ordinary PDF editing can invalidate digital signatures. If it must remain an accessible document, assess the effect of every conversion on tags, reading order and alternative text.

The best workflow is therefore conditional: identify the page types, understand the delivery requirement, try the least destructive change, inspect the output, and only then make a stronger trade-off. File size is one acceptance criterion. Readability, completeness, accessibility and document behaviour are separate criteria that deserve their own checks.

For scanned text that needs search, continue with OCR PDF. For a packet containing unnecessary pages, use Remove pages on a copy. For the distinction between a visual page and selectable words, read Why PDF text copies in the wrong order.