API reference · Context.dev

context-dev/extract-structured-data

Integrate this model through SandBase's unified API, with production-ready schemas and examples.

Production endpoint

Send your first request

OpenAI-compatible endpoint with unified authentication and usage tracking.

POSThttps://api.sandbase.ai/v1/run
Model IDcontext-dev/extract-structured-data
01

Input Schema

13 parameters · 2 required · 11 optional

ParameterTypeRequiredDescription
urlstringRequiredStarting website URL to crawl and extract from.
schemaobjectRequiredJSON Schema describing the object to return. Generate it from a Zod or Pydantic model, or hand-write it — Context.dev fills exactly this shape.
maxAgeMsintegerOptionalReuse a cached result younger than this many milliseconds. Default 86400000 (1 day), max 2592000000 (30 days). Set 0 to always fetch fresh. · Min: 0 · Max: 2592000000
maxDepthintegerOptionalMaximum link depth from the starting URL (0 = only the starting page). Unlimited when omitted. · Min: 0 · Max: 9007199254740991
maxPagesintegerOptionalMaximum number of pages to analyze. Default 5, hard cap 50. Does NOT change the price — extraction is billed per call. · Min: 1 · Max: 50
factCheckbooleanOptionalWhen true, every returned value must be grounded in text stated on the page and unsupported fields come back null/empty. When false (default), reasonable inferences are allowed while verifiable specifics stay faithful to the source.
timeoutMSintegerOptionalUpstream timeout in milliseconds (max 300000). Context.dev aborts the call with a 408 when exceeded. Keep it below the endpoint's requestTimeoutMs so the provider answers before the platform's own budget expires. · Min: 1000 · Max: 300000
waitForMsintegerOptionalExtra browser wait in milliseconds after page load before the content is captured (0-30000). Useful for JavaScript-heavy pages. · Min: 0 · Max: 30000
stopAfterMsintegerOptionalSoft time budget for the crawl phase in milliseconds (10000-110000, default 80000). · Min: 10000 · Max: 110000
instructionsstringOptionalExtraction guidance: which facts to prioritize, how to interpret ambiguous fields. · Max length: 2000
includeFramesbooleanOptionalInclude iframe contents in the Markdown handed to the extractor. Default false.
followSubdomainsbooleanOptionalFollow links on subdomains of the starting URL's domain. Default false.
settleAnimationsbooleanOptionalWait briefly for animations to settle before each page is read. Default false.
02

Output Schema

FieldTypeDescription
idstringUnique identifier for the generation task
statusstringTask status: pending, running, completed, failed, timeout
modelstringModel used for the generation
outputsarrayArray of output items
outputs[].urlstringURL of the generated artifact
outputs[].content_typestringMIME type (e.g. image/png, video/mp4)
errorobject | nullError details if failed, null on success
error.typestringMachine-readable error type code
error.messagestringHuman-readable error description

Async Workflow

This model uses asynchronous execution. Submit a request and poll for the result.

  1. Submit — POST to /v1/run, receive an id
  2. Poll — GET /v1/run/{id} until status is completed, failed, or timeout
  3. Retrieve — Read outputs from the completed response
03

Code Examples

Ready-to-run snippets

# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
  "model": "context-dev/extract-structured-data",
  "prompt": "a beautiful sunset over mountains"
}'

# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
  -H "Authorization: Bearer YOUR_API_KEY"