Home Blog How to Extract a Table from PDF into Excel or CSV

How to Extract a Table from PDF into Excel or CSV

Related to this article: PDF to Excel — Convert OnlineOpen converter →

A report, a statement, or a price list often exists only as a PDF, yet you need to work with the data as a regular spreadsheet — sorting, calculating, filtering. For that, the table needs to be extracted from the PDF into Excel or CSV.

Why this is harder than a regular conversion

Unlike XLSX or CSV, a PDF doesn't store a table's structure — only how the page looks visually, meaning where each piece of text sits. So extracting a table works differently from a regular format conversion: the text is first extracted while preserving its original layout on the page, and then columns are identified heuristically, based on gaps of two or more spaces between values.

This approach works well for simple tables with clearly defined column boundaries — reports or price lists with a neat grid, for example. It handles complex layouts worse: merged cells, multi-line values within a single cell, or unusual alignment, where columns can end up misaligned.

How to extract a table into Excel

The PDF to Excel converter lays the extracted data out into XLSX cells — the result is immediately ready for sorting, formula calculations, and formatting like a normal spreadsheet. Upload the PDF and download the resulting XLSX.

How to extract a table into CSV

If you don't need a full Excel file but a plain text format for importing into a database or another program, use the PDF to CSV converter — the same extraction principle, but the result is saved as plain delimited text.

Why it won't work on a scanned PDF

If the PDF is a scan or a photograph of a document, the file physically has no text layer, so there's nothing to extract — the tool will just see an image. It's easy to check: try selecting text in the PDF with your mouse. If it selects, there's a text layer and the conversion will work. If the whole page gets selected as one block, like an image, you're looking at a scan, and it needs text recognition (OCR) first.

Bottom line

Extracting a table from PDF works well for simple tables with a text layer and clear columns, and noticeably worse for complex layouts and scans. It's worth checking whether the text in the file can be selected before converting, so you know what to expect from the result.

← All articles