Extract

Get the status and result of an async extract job

GET
/api/v1/extract/{job_id}

Get the status and result of an async extract job

Poll for the status and result of an async extract job submitted via POST /api/v1/extract with options.async=true. Returns the same envelope shape as the synchronous extract endpoint once status is completed.

Pagination: when completed, result.data is paginated with a fixed page size of 15. Use the page query param (1-based) to navigate; result.pagination reports total_items, total_pages, has_next, has_prev.

Recommended polling cadence: 1s for the first 10s, then 5s, capped at 30s. Stop polling once status is in {completed, failed}.

Authorization

bearerAuth
AuthorizationBearer <token>

Console session token (Authorization: Bearer <session>) or product API key (Authorization: Bearer <api-key> or x-api-key). Session tokens are validated via Better Auth get-session; API keys against the shared database.

In: header

Path Parameters

job_id*Job Id

Query Parameters

page?|

1-based page index for navigating result.data. Fixed page size of 15 items. Returns 400 if the page is out of range.

Header Parameters

authorization?string|null
x-api-key?string|null

Response Body

application/json

application/json

application/json

curl -X GET "https://example.com/api/v1/extract/string"
{  "id": "string",  "status": "string",  "created_at": "2019-08-24T14:15:22Z",  "completed_at": "2019-08-24T14:15:22Z",  "processing_time_ms": 0,  "document": {    "filename": "string",    "page_count": 0,    "file_size_bytes": 0,    "mime_type": "string"  },  "result": {    "data": [      {}    ],    "images": [      {        "index": 0,        "page": 0,        "url": "string",        "key": "string",        "mimetype": "string",        "width": 0,        "height": 0      }    ],    "pagination": {      "page": 0,      "page_size": 0,      "total_items": 0,      "total_pages": 0,      "has_next": true,      "has_prev": true    }  },  "usage": {    "pages_processed": 0,    "images_uploaded": 0  },  "progress": {    "percentage": 0,    "pages_processed": 0  }}

Extract structured data from a document POST

Extract structured data from a document Pull specific fields from a document into a typed schema. Accepts either a **file upload** (multipart/form-data) or a **document URL** (JSON body), plus a **JSON Schema** (the `schema` field) describing what to extract. ### Sync mode (default) Blocks until extraction completes and returns **200** with the full result. ```bash curl -X POST https://api.context212.com/api/v1/extract \ -H 'Authorization: Bearer $TOKEN' \ -F file=@invoice.pdf \ -F 'schema={"type":"object","properties":{"invoice_number":{"type":"string"}}}' ``` ### Scope (`options.scope`) - **`page`** (default): apply the schema independently to each page. `result.data` has one object per page (fields missing on a page are `null`). - **`document`**: apply the schema once to the full document (all pages joined). `result.data` is a single-element array with the consolidated object. Use this when fields and arrays span multiple pages (e.g. an auction notice with lots). ### Async mode (`options.async = true`) Not implemented yet — returns **501**. ### Embedded images (`options.include_images`, **default true** for extract) Extract uploads every embedded PDF image, injects `![image N](url)` into the parsed markdown **before** the LLM runs (so schema fields can capture URLs), and returns the same list on `result.images`. Pass `options.include_images: false` to skip uploads. For multipart: `-F 'options={"scope":"document"}'` (images on by default). **Supported file types:** `.pdf`, `.png`, `.jpg`, `.jpeg`, `.pptx`, `.docx`, `.xlsx`, `.html`, `.xhtml` **Sync limits:** 20 MB file size, 15 pages.