Parse

Render a document as PDF

POST
/api/v1/preview

Render a document as PDF

Convert a document into a PDF suitable for inline preview.

Accepts a file upload (multipart/form-data), a document URL, or a file_id referencing a document already ingested in the platform (JSON body). The conversion is synchronous — the PDF bytes are returned directly.

Supported file types: .pdf, .png, .jpg, .jpeg, .pptx, .ppt, .odp, .docx, .odt, .doc, .html, .xhtml

Size limit: 20 MB.

Authorization

bearerAuth
AuthorizationBearer <token>

Console session token (Authorization: Bearer <session>) or product API key (Authorization: Bearer <api-key> or x-api-key). Session tokens are validated via Better Auth get-session; API keys against the shared database.

In: header

Header Parameters

authorization?string|null
x-api-key?string|null

Request Body

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/pdf

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/api/v1/preview" \  -H "Content-Type: application/json" \  -d '{    "document": "https://example.com/report.docx"  }'
"string"

Get the status and result of an async parse job GET

Get the status and result of an async parse job Poll for the status and result of an async parse job submitted via `POST /api/v1/parse` with `options.async=true`. Returns the same envelope shape as the synchronous parse endpoint once `status` is `completed`. ```bash curl https://api.context212.com/api/v1/parse/parse_Kg \ -H 'Authorization: Bearer $TOKEN' ``` **Recommended polling cadence:** 1s for the first 10s, then 5s, capped at 30s. Stop polling once `status` is in `{completed, failed}`.

Extract structured data from a document POST

Extract structured data from a document Pull specific fields from a document into a typed schema. Accepts either a **file upload** (multipart/form-data) or a **document URL** (JSON body), plus a **JSON Schema** (the `schema` field) describing what to extract. ### Sync mode (default) Blocks until extraction completes and returns **200** with the full result. ```bash curl -X POST https://api.context212.com/api/v1/extract \ -H 'Authorization: Bearer $TOKEN' \ -F file=@invoice.pdf \ -F 'schema={"type":"object","properties":{"invoice_number":{"type":"string"}}}' ``` ### Scope (`options.scope`) - **`page`** (default): apply the schema independently to each page. `result.data` has one object per page (fields missing on a page are `null`). - **`document`**: apply the schema once to the full document (all pages joined). `result.data` is a single-element array with the consolidated object. Use this when fields and arrays span multiple pages (e.g. an auction notice with lots). ### Async mode (`options.async = true`) Not implemented yet — returns **501**. ### Embedded images (`options.include_images`, **default true** for extract) Extract uploads every embedded PDF image, injects `![image N](url)` into the parsed markdown **before** the LLM runs (so schema fields can capture URLs), and returns the same list on `result.images`. Pass `options.include_images: false` to skip uploads. For multipart: `-F 'options={"scope":"document"}'` (images on by default). **Supported file types:** `.pdf`, `.png`, `.jpg`, `.jpeg`, `.pptx`, `.docx`, `.xlsx`, `.html`, `.xhtml` **Sync limits:** 20 MB file size, 15 pages.