Guide · September 2026

Extract tables from PDF

A PDF table is not a table — it's coordinates. Getting it out cleanly is where most PDF pipelines fall over. This page walks through the common failure modes (side-by-side layouts, multi-page tables, EU number formats) and shows how the Docule API handles them.

Why PDF table extraction is hard

PDF has no table primitive. What looks like a table to a reader is really a set of text runs at absolute (x, y) coordinates, sometimes with a few horizontal rules drawn behind them. A parser has to reconstruct the row/column grid from geometry alone, and then work out which column each cell belongs to.

The most common failures we see:

How Docule handles them

Docule's pipeline starts with per-page geometry, layers a table detector, then applies a set of financial-document-specific fixups. The interesting ones:

What a clean extraction looks like

Input page: a cash flow statement with two columns per year (2026 and 2025), 12 rows each, in MEUR.

Output JSON (trimmed):

{
  "type": "table",
  "columns": ["item", "2026", "2025"],
  "value": [
    ["Net cash from operating activities", "3921", "3204"],
    ["Net cash used in investing activities", "-1150", "-980"],
    ["Net cash from financing activities", "-402", "-317"],
    ["Cash and cash equivalents, end of period", "8842", "6473"]
  ],
  "unit": "MEUR",
  "currency": "EUR",
  "bbox": [72, 210, 540, 620],
  "value_normalized": [
    {"item": "Net cash from operating activities", "2026": 3.921e9, "2025": 3.204e9},
    {"item": "Net cash used in investing activities", "2026": -1.150e9, "2025": -0.980e9},
    {"item": "Net cash from financing activities", "2026": -0.402e9, "2025": -0.317e9},
    {"item": "Cash and cash equivalents, end of period", "2026": 8.842e9, "2025": 6.473e9}
  ]
}

Minimal end-to-end

import os, time, requests

API = "https://docule.dev/api/v1"
KEY = os.environ["DOCULE_API_KEY"]

job = requests.post(
    f"{API}/parse",
    headers={"X-API-Key": KEY},
    files={"file": open("annual_report.pdf", "rb")},
    params={"formats": "json", "merge_continuations": "true"},
).json()

# Poll until done (elided — see the PDF-to-JSON guide)

result = requests.get(
    f"{API}/result/{job['job_id']}",
    headers={"X-API-Key": KEY},
    params={"formats": "json"},
).json()

for page in result["result"]["pages"]:
    for item in page["items"]:
        if item["type"] == "table":
            print(f"Page {page['page']} — {len(item['value'])} rows, cols: {item['columns']}")

How Docule compares to other approaches

ApproachSimple tablesSide-by-side / multi-columnEU units + locales
Copy-paste from PreviewOKFailsFails
pdftotext CLIOKFailsFails
Camelot / Tabula (Python)GoodSometimesFails
Generic AI parserGoodSometimesSometimes
DoculeGoodGoodGood
Try it on your worst PDF first. The right question is not "does it handle simple tables" — it's "does it handle the messy statement that broke your last pipeline". pdftotext.cc lets you drop one PDF and see the raw extraction quality with no signup.

See also

Get your API key

Free tier is 6,000 credits per month — enough to parse ~100 pages of tables. No card required.

Get API Key →