Extract

Extract structured data from a document

POST
/api/v1/extract

Extract structured data from a document

Pull specific fields from a document into a typed schema.

Accepts either a file upload (multipart/form-data) or a document URL (JSON body), plus a JSON Schema (the schema field) describing what to extract.

Sync mode (default)

Blocks until extraction completes and returns 200 with the full result.

curl -X POST https://api.context212.com/api/v1/extract \  -H 'Authorization: Bearer $TOKEN' \  -F file=@invoice.pdf \  -F 'schema={"type":"object","properties":{"invoice_number":{"type":"string"}}}'

Async mode (options.async = true)

Not implemented yet — returns 501.

For multipart uploads, pass options as a JSON-encoded form field: -F 'options={"async":true}'.

Supported file types: .pdf, .png, .jpg, .jpeg, .pptx, .docx, .xlsx, .html, .xhtml

Sync limits: 20 MB file size, 15 pages.

Authorization

bearerAuth
AuthorizationBearer <token>

Console session token (Authorization: Bearer <session>) or product API key (Authorization: Bearer <api-key> or x-api-key). Session tokens are validated via Better Auth get-session; API keys against the shared database.

In: header

Header Parameters

authorization?string|null
x-api-key?string|null

Request Body

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/api/v1/extract" \  -H "Content-Type: application/json" \  -d '{    "schema": {}  }'
{  "id": "string",  "status": "string",  "created_at": "2019-08-24T14:15:22Z",  "completed_at": "2019-08-24T14:15:22Z",  "processing_time_ms": 0,  "document": {    "filename": "string",    "page_count": 0,    "file_size_bytes": 0,    "mime_type": "string"  },  "result": {    "data": [      {}    ],    "pagination": {      "page": 0,      "page_size": 0,      "total_items": 0,      "total_pages": 0,      "has_next": true,      "has_prev": true    }  },  "usage": {    "pages_processed": 0  },  "progress": {    "percentage": 0,    "pages_processed": 0  }}