PDF to Excel
Move document text into a spreadsheet.
Infer consistent columns from aligned text and export editable cells with number detection. Check merged cells and complex tables. Original spreadsheet formulas cannot be recovered from PDF.
How to use PDF to Excel
- Open your document from your device.
- Adjust the settings and preview the result.
- Read text results on screen or download your file. Your original remains untouched.
Local by design
Your documents are processed on your device. No account, payment or document upload is required.
Enable JavaScript to use the document workspace.
Capabilities and limitationsExtract table values into cells you can review
PDF to Excel is useful when aligned text in a PDF needs to become a starting dataset. SandboxPDF infers columns from positions and creates editable cells with basic number detection. A visual table is not necessarily stored as real rows and columns, so the resulting workbook requires review before calculations or import into another system.
Begin by identifying the table and its conventions. Check whether the source is a scan, whether values use decimal commas or points, and whether headings repeat across pages. Keep the original PDF and a raw extraction copy. That gives you a reference when correcting ambiguous cells and helps preserve a traceable path to the final dataset.
Convert and inspect a representative section
- Open the PDF and choose the relevant pages.
- Select aligned-column detection for table-oriented output.
- Create the XLSX and open it in your spreadsheet application.
- Compare row count, column labels and important values with the source.
- Save a separately reviewed working copy before adding analysis formulas.
For image-only tables, OCR is required before meaningful text can be extracted. Recognition can confuse digits and punctuation, while column detection can place correct text in the wrong cell. These are separate possible errors. A neat-looking sheet does not prove that either stage was accurate.
Verify types and relationships
Check identifiers with leading zeros, long reference numbers and codes that resemble dates. These may need to remain text rather than numeric values. Inspect negative amounts, percentages and currency units. Automatic number detection is a convenience, not a guarantee that every displayed string has the intended meaning in a spreadsheet.
Wrapped descriptions can create extra rows or become detached from their values. Merged headings can shift inferred columns. Check the transition between source pages, where repeated headers and subtotals are common. Trace several complete records across all columns instead of checking isolated numbers in one column.
PDFs do not normally contain the original spreadsheet formulas. An extracted total is a displayed value, not a recovered calculation. Rebuild formulas explicitly if you need a working model. Use recalculated totals as one check, but remember that offsetting errors can produce a matching sum even when individual records are wrong.
Prepare a usable final workbook
Give columns clear names and decide how blanks, zeros and “not applicable” values should be represented. Remove repeated headers only after confirming they are not data. Record important manual corrections and retain enough source references for another person to verify them. A cleaned dataset should be understandable without guessing how it was extracted.
If the source uses a complex table structure, consider requesting a CSV or XLSX from its publisher. A structured original can preserve relationships that a PDF presentation discarded. For a small critical table, careful manual transcription with an independent check may be more reliable than repairing a large uncertain extraction.
Open the saved workbook again to confirm that types and formulas behave as intended in the destination application. Read the table verification guide for detailed checks and the OCR guide for scanned inputs. The aim is trustworthy reusable data, not merely a file with the right extension.
Frequently asked questions
Are the original formulas recovered?
No. The PDF normally contains displayed values, not the spreadsheet calculation history.
Why are some values in the wrong column?
Columns are inferred from text placement. Wrapped, merged or irregular layouts can need manual correction.
Can I convert scanned tables directly?
A usable text layer is needed. Apply and verify OCR first, then check the extracted cell relationships.
Should every digit-only cell be a number?
No. Identifiers and codes can require text storage to preserve leading zeros and exact characters.
Is a matching total enough to prove accuracy?
No. Check individual records, types and source relationships as well as aggregate totals.