Guide · September 2026

PDF to JSON — a working guide

Converting a PDF to structured JSON so an LLM, a database, or a downstream job can consume it. This page walks you from a raw PDF to a JSON payload with page structure, tables, and bounding-box anchors, using a REST call you can paste into any language.

What you get

Docule returns per-page JSON that keeps the document's structure — headings, paragraphs, tables, list items — as typed items. Every item has a bbox anchor, so you can highlight the source region back on the original PDF.

A single-page response looks like this:

{
  "job_id": "job_9f8e7d",
  "status": "completed",
  "result": {
    "pages": [
      {
        "page": 1,
        "text": "Net sales grew 8.4% in Q2, driven by...",
        "markdown": "# Net sales grew 8.4% in Q2...",
        "items": [
          {
            "type": "heading",
            "value": "Q2 2026 Financial Summary",
            "level": 1,
            "bbox": [72, 68, 540, 92]
          },
          {
            "type": "paragraph",
            "value": "Net sales grew 8.4% in Q2...",
            "bbox": [72, 108, 540, 168]
          },
          {
            "type": "table",
            "value": [
              ["", "Q2 2026", "Q2 2025"],
              ["Net sales", "25758", "23759"],
              ["Operating profit", "3921", "3204"]
            ],
            "columns": ["metric", "q2_2026", "q2_2025"],
            "bbox": [72, 200, 540, 340]
          }
        ]
      }
    ]
  }
}

Step 1 — Get an API key

Sign up at docule.dev. The free plan is 6,000 credits (about 100 pages) per month with no credit card. Copy the key from the dashboard; it looks like docule_live_abc123....

Step 2 — POST the PDF

The parse endpoint accepts a multipart upload and returns a job_id. Parsing is asynchronous — you'll poll for completion in step 3.

# curl
curl -X POST https://docule.dev/api/v1/parse \
  -H "X-API-Key: $DOCULE_API_KEY" \
  -F "file=@report.pdf" \
  -G --data-urlencode "formats=json"
# → 202 Accepted, { "job_id": "job_9f8e7d" }

Step 3 — Poll for completion

Parse jobs typically finish in under 5 seconds per page. Poll the status endpoint every 1–2 seconds:

curl https://docule.dev/api/v1/status/job_9f8e7d \
  -H "X-API-Key: $DOCULE_API_KEY"
# → { "status": "completed", "progress": 1.0 }

Step 4 — Fetch the JSON

curl "https://docule.dev/api/v1/result/job_9f8e7d?formats=json" \
  -H "X-API-Key: $DOCULE_API_KEY"

The full Python example

End-to-end, no dependencies beyond the standard requests library:

import os, time, requests

API = "https://docule.dev/api/v1"
KEY = os.environ["DOCULE_API_KEY"]

# 1. Submit
job = requests.post(
    f"{API}/parse",
    headers={"X-API-Key": KEY},
    files={"file": open("report.pdf", "rb")},
    params={"formats": "json"},
).json()

job_id = job["job_id"]

# 2. Poll
while True:
    status = requests.get(
        f"{API}/status/{job_id}",
        headers={"X-API-Key": KEY},
    ).json()
    if status["status"] == "completed":
        break
    if status["status"] == "failed":
        raise RuntimeError(status.get("error", "parse failed"))
    time.sleep(1.5)

# 3. Fetch the JSON
result = requests.get(
    f"{API}/result/{job_id}",
    headers={"X-API-Key": KEY},
    params={"formats": "json"},
).json()

for page in result["result"]["pages"]:
    for item in page["items"]:
        if item["type"] == "table":
            print(f"Page {page['page']}: {len(item['value'])} rows")

Response schema, in short

Every response has the same shape. The interesting fields:

Why bounding boxes matter. When an LLM answers a question over the parsed JSON, you can point back to the exact rectangle on the PDF. That is how you build a "show me the source" button that a compliance team will actually trust.

Common pitfalls

Scanned PDFs. Text-based PDFs are fast (~1s/page). Scans need OCR and cost more credits — Docule detects this and escalates automatically, but a 200-page scanned annual report will take a minute or two.

Multi-column layouts. Newspapers, brochures and some annual reports use 2- or 3-column layouts. Docule reconstructs reading order — no code changes on your side.

Tables that span pages. Set merge_continuations=true in the query string and Docule joins split tables across pages before returning.

Non-English documents. Nordic and EU locale (comma decimals, MEUR/MSEK/kEUR units, dmy dates) is handled by default. Set locale=fi or locale=sv to force detection when a document has mixed content.

Related

Get your API key

6,000 credits per month, no credit card required. Parse your first PDF in under a minute.

Get API Key →