How to get a table out of a PDF and into Excel
Copying a table out of a PDF by hand is where afternoons go to die — and where the typos that break a total come from. It can be done properly, but it is worth knowing what the tool is guessing.
How a table is found
A PDF does not contain tables. It contains text at coordinates and lines drawn on the page. When the table has ruling lines, a converter can rebuild the grid exactly: the vertical lines are the columns and the horizontal ones are the rows.
When there are no lines, the only clue left is alignment — fragments that start at the same x position, row after row. That works well and it is still a deduction, which is why a serious tool marks those tables as detected rather than certain.
Two things are often mishandled and worth checking in the result: a header that spans several columns (a merged cell) and a table that continues on the next page with its header repeated. Both should end up as one table, not as three.
Numbers that must arrive as numbers
A figure that lands in Excel as text cannot be added up, sorted or charted, and Excel marks it with a little green triangle. Getting this right is most of the value of the conversion.
The hard part is that the same string means different things in different countries: 1.240 is one thousand two hundred and forty in Spain and one point two four in the United States. A careful converter uses the clues in the document — the other figures in the same column, the currency symbol, the decimals — and when there is genuinely no clue, it keeps the value as text and tells you, instead of silently turning your total into something a thousand times wrong.
The same applies to dates: 03/04/2026 is the third of April or the fourth of March depending on where the document comes from. Watch that column closely after converting.
Check it in one minute
Select the column of amounts. Excel's status bar shows a sum only if the cells are numbers — if there is no sum, they arrived as text.
Compare the last row against the PDF: totals are where an extra or missing row shows up immediately.
Check one date and one percentage. If those two are right, the rest of the column almost always is.
What about a scanned table
If the pages are photographs, there is no text to extract: reading the words out of the image needs OCR, which CLEKARA does not do yet — it says so rather than handing you an empty sheet.
Whatever the tool, a crooked or dark scan reads badly. Scanning again straight, at 300 DPI and in good light changes the result more than any setting.
You can try your own file in PDF to Excel — it runs in your browser, so the document stays on your computer.