Extract tables from PDF
A PDF table is not a table — it's coordinates. Getting it out cleanly is where most PDF pipelines fall over. This page walks through the common failure modes (side-by-side layouts, multi-page tables, EU number formats) and shows how the Docule API handles them.
Why PDF table extraction is hard
PDF has no table primitive. What looks like a table to a reader is really a set of text runs at absolute (x, y) coordinates, sometimes with a few horizontal rules drawn behind them. A parser has to reconstruct the row/column grid from geometry alone, and then work out which column each cell belongs to.
The most common failures we see:
- Side-by-side blocks read as one wide table. A cash flow statement often shows current year and prior year in two columns, side by side. A generic parser sees a 10-column wide grid; the correct answer is two 5-column tables.
- Header row merged with the first data row. When the header line is close to the first row, span reconstruction pulls them into one line.
- Split tables across pages. A 40-row income statement split at page 3 → 4 becomes two half-tables unless the parser stitches them back.
- Unit prefixes in the header (MEUR, MSEK, kEUR). Extractors return raw digits without applying the multiplier.
25758 MEURgets stored as25758, and the finance team lands on the wrong billion. - Comma decimals. Nordic and EU PDFs write
1 234,5. A parser tuned for US format reads it as1234or truncates to1. - Empty header column for the row label. A statement's first column ("Net sales", "Cost of goods sold", …) usually has a blank header. Naive extractors give it a placeholder like
Unnamed: 0.
How Docule handles them
Docule's pipeline starts with per-page geometry, layers a table detector, then applies a set of financial-document-specific fixups. The interesting ones:
- Side-by-side detection. When the column geometry shows two clean sub-grids separated by a visual gutter, they are emitted as two separate tables sharing the same header row schema.
- Header repair. Ambiguous header rows are validated against the following data rows (does the header column align with the data column?). Broken headers are re-inferred from context — page text, section titles, common financial statement shapes.
- Multi-page merge. Set
merge_continuations=trueand Docule stitches split tables back together, using the row-label column and the header schema as the join key. - Unit and locale normalisation. The
value_normalizedfield on each row applies the header's unit multiplier (MEUR × 1e6), converts comma-decimals to dot-decimals, and returns floats you can consume directly. - Row-sum validation. For income statements and balance sheets, sub-totals are checked against their components. Rows that fail are flagged in the
validationmetadata rather than silently returned as-is.
What a clean extraction looks like
Input page: a cash flow statement with two columns per year (2026 and 2025), 12 rows each, in MEUR.
Output JSON (trimmed):
{
"type": "table",
"columns": ["item", "2026", "2025"],
"value": [
["Net cash from operating activities", "3921", "3204"],
["Net cash used in investing activities", "-1150", "-980"],
["Net cash from financing activities", "-402", "-317"],
["Cash and cash equivalents, end of period", "8842", "6473"]
],
"unit": "MEUR",
"currency": "EUR",
"bbox": [72, 210, 540, 620],
"value_normalized": [
{"item": "Net cash from operating activities", "2026": 3.921e9, "2025": 3.204e9},
{"item": "Net cash used in investing activities", "2026": -1.150e9, "2025": -0.980e9},
{"item": "Net cash from financing activities", "2026": -0.402e9, "2025": -0.317e9},
{"item": "Cash and cash equivalents, end of period", "2026": 8.842e9, "2025": 6.473e9}
]
}
Minimal end-to-end
import os, time, requests
API = "https://docule.dev/api/v1"
KEY = os.environ["DOCULE_API_KEY"]
job = requests.post(
f"{API}/parse",
headers={"X-API-Key": KEY},
files={"file": open("annual_report.pdf", "rb")},
params={"formats": "json", "merge_continuations": "true"},
).json()
# Poll until done (elided — see the PDF-to-JSON guide)
result = requests.get(
f"{API}/result/{job['job_id']}",
headers={"X-API-Key": KEY},
params={"formats": "json"},
).json()
for page in result["result"]["pages"]:
for item in page["items"]:
if item["type"] == "table":
print(f"Page {page['page']} — {len(item['value'])} rows, cols: {item['columns']}")
How Docule compares to other approaches
| Approach | Simple tables | Side-by-side / multi-column | EU units + locales |
|---|---|---|---|
| Copy-paste from Preview | OK | Fails | Fails |
pdftotext CLI | OK | Fails | Fails |
| Camelot / Tabula (Python) | Good | Sometimes | Fails |
| Generic AI parser | Good | Sometimes | Sometimes |
| Docule | Good | Good | Good |
See also
- Guide: PDF to JSON — the request/response walkthrough
- Docule vs LlamaParse
- Docule vs Unstructured
- Full API reference
Get your API key
Free tier is 6,000 credits per month — enough to parse ~100 pages of tables. No card required.
Get API Key →