PDF to JSON — a working guide
Converting a PDF to structured JSON so an LLM, a database, or a downstream job can consume it. This page walks you from a raw PDF to a JSON payload with page structure, tables, and bounding-box anchors, using a REST call you can paste into any language.
What you get
Docule returns per-page JSON that keeps the document's structure — headings, paragraphs, tables, list items — as typed items. Every item has a bbox anchor, so you can highlight the source region back on the original PDF.
A single-page response looks like this:
{
"job_id": "job_9f8e7d",
"status": "completed",
"result": {
"pages": [
{
"page": 1,
"text": "Net sales grew 8.4% in Q2, driven by...",
"markdown": "# Net sales grew 8.4% in Q2...",
"items": [
{
"type": "heading",
"value": "Q2 2026 Financial Summary",
"level": 1,
"bbox": [72, 68, 540, 92]
},
{
"type": "paragraph",
"value": "Net sales grew 8.4% in Q2...",
"bbox": [72, 108, 540, 168]
},
{
"type": "table",
"value": [
["", "Q2 2026", "Q2 2025"],
["Net sales", "25758", "23759"],
["Operating profit", "3921", "3204"]
],
"columns": ["metric", "q2_2026", "q2_2025"],
"bbox": [72, 200, 540, 340]
}
]
}
]
}
}
Step 1 — Get an API key
Sign up at docule.dev. The free plan is 6,000 credits (about 100 pages) per month with no credit card. Copy the key from the dashboard; it looks like docule_live_abc123....
Step 2 — POST the PDF
The parse endpoint accepts a multipart upload and returns a job_id. Parsing is asynchronous — you'll poll for completion in step 3.
# curl
curl -X POST https://docule.dev/api/v1/parse \
-H "X-API-Key: $DOCULE_API_KEY" \
-F "file=@report.pdf" \
-G --data-urlencode "formats=json"
# → 202 Accepted, { "job_id": "job_9f8e7d" }
Step 3 — Poll for completion
Parse jobs typically finish in under 5 seconds per page. Poll the status endpoint every 1–2 seconds:
curl https://docule.dev/api/v1/status/job_9f8e7d \
-H "X-API-Key: $DOCULE_API_KEY"
# → { "status": "completed", "progress": 1.0 }
Step 4 — Fetch the JSON
curl "https://docule.dev/api/v1/result/job_9f8e7d?formats=json" \
-H "X-API-Key: $DOCULE_API_KEY"
The full Python example
End-to-end, no dependencies beyond the standard requests library:
import os, time, requests
API = "https://docule.dev/api/v1"
KEY = os.environ["DOCULE_API_KEY"]
# 1. Submit
job = requests.post(
f"{API}/parse",
headers={"X-API-Key": KEY},
files={"file": open("report.pdf", "rb")},
params={"formats": "json"},
).json()
job_id = job["job_id"]
# 2. Poll
while True:
status = requests.get(
f"{API}/status/{job_id}",
headers={"X-API-Key": KEY},
).json()
if status["status"] == "completed":
break
if status["status"] == "failed":
raise RuntimeError(status.get("error", "parse failed"))
time.sleep(1.5)
# 3. Fetch the JSON
result = requests.get(
f"{API}/result/{job_id}",
headers={"X-API-Key": KEY},
params={"formats": "json"},
).json()
for page in result["result"]["pages"]:
for item in page["items"]:
if item["type"] == "table":
print(f"Page {page['page']}: {len(item['value'])} rows")
Response schema, in short
Every response has the same shape. The interesting fields:
result.pages[]— one entry per source page.page.text— plain-text render of the page.page.markdown— Markdown render, useful as the RAG chunk body.page.items[]— structured elements. Each hastype(heading,paragraph,table,list,image, …),value, andbbox.bbox—[x0, y0, x1, y1]in PDF-user-space coordinates. Combine with the source PDF's page height/width to draw a highlight.page.key_figures— deterministic prose key-figures (revenue, EBIT, EPS, equity ratio, …) pulled from press-release-style sentences. Each carriesmetric,value,value_base,currency,scale,comparison(prior-period),period,confidence, and the verbatimsourcesentence. Runs on every page — no extra request parameter — and skips primary financial statements so it never double-reports what the table pipeline already emitted.
Common pitfalls
Scanned PDFs. Text-based PDFs are fast (~1s/page). Scans need OCR and cost more credits — Docule detects this and escalates automatically, but a 200-page scanned annual report will take a minute or two.
Multi-column layouts. Newspapers, brochures and some annual reports use 2- or 3-column layouts. Docule reconstructs reading order — no code changes on your side.
Tables that span pages. Set merge_continuations=true in the query string and Docule joins split tables across pages before returning.
Non-English documents. Nordic and EU locale (comma decimals, MEUR/MSEK/kEUR units, dmy dates) is handled by default. Set locale=fi or locale=sv to force detection when a document has mixed content.
Related
- Guide: Extract tables from PDF — the tricky case
- Docule vs LlamaParse
- Docule vs Unstructured
- Full API reference
Get your API key
6,000 credits per month, no credit card required. Parse your first PDF in under a minute.
Get API Key →