API reference · Context.dev
context-dev/crawl-site
Integrate this model through SandBase's unified API, with production-ready schemas and examples.
Production endpoint
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
POST
https://api.sandbase.ai/v1/runModel ID
context-dev/crawl-site01
Input Schema
20 parameters · 1 required · 19 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
url | string | Required | Starting URL for the crawl, including the http:// or https:// scheme. |
pdf | object | Optional | PDF parsing controls. |
zdr | string | Optional | Set 'enabled' to bypass Context.dev's shared caches and omit request and response content from its retained usage logs. Requires zero data retention on the Context.dev organization, otherwise the call fails with ZDR_NOT_ENABLED. · Options: enabled, disabled enableddisabled |
country | string | Optional | Two-letter ISO 3166-1 alpha-2 country code (e.g. 'us', 'gb', 'de'). Must be one of Context.dev's supported countries. · Min length: 2 · Max length: 2 |
maxAgeMs | integer | Optional | Reuse a cached result younger than this many milliseconds. Default 86400000 (1 day), max 2592000000 (30 days). Set 0 to always fetch fresh. · Min: 0 · Max: 2592000000 |
maxDepth | integer | Optional | Maximum link depth from the starting URL (0 = only the starting page). Unlimited when omitted. · Min: 0 · Max: 9007199254740991 |
maxPages | integer | Optional | Maximum number of pages to crawl. Default 100, hard cap 500. This is also what sizes the admission hold — one credit per page. · Min: 1 · Max: 500 |
urlRegex | string | Optional | Only URLs matching this pattern are followed and scraped. |
timeoutMS | integer | Optional | Upstream timeout in milliseconds (max 300000). Context.dev aborts the call with a 408 when exceeded. Keep it below the endpoint's requestTimeoutMs so the provider answers before the platform's own budget expires. · Min: 1000 · Max: 300000 |
waitForMs | integer | Optional | Extra browser wait in milliseconds after page load before the content is captured (0-30000). Useful for JavaScript-heavy pages. · Min: 0 · Max: 30000 |
stopAfterMs | integer | Optional | Soft time budget for the whole crawl in milliseconds (10000-110000, default 80000). When exceeded, the pages collected so far are returned instead of continuing. · Min: 10000 · Max: 110000 |
includeLinks | boolean | Optional | Preserve hyperlinks in the Markdown output. Default true. |
includeFrames | boolean | Optional | Render iframe contents into the output. Default false. |
includeImages | boolean | Optional | Include image references in the Markdown output. Default false. |
excludeSelectors | string[] | Optional | CSS selectors to remove before conversion. Applied after includeSelectors; exclusion wins on a double match. |
followSubdomains | boolean | Optional | Follow links on subdomains of the starting URL's domain (e.g. docs.example.com from example.com). www and apex are always equivalent. Default false. |
includeSelectors | string[] | Optional | CSS selectors. When provided, only matching subtrees are kept before each page is converted to Markdown. |
settleAnimations | boolean | Optional | Wait briefly for animations to settle before capturing each page. Default false. |
useMainContentOnly | boolean | Optional | Keep only the page's main content, dropping headers, footers, sidebars, and navigation when detectable. Default false. |
shortenBase64Images | boolean | Optional | Truncate inline base64 image payloads to keep the output small. Default true. |
02
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
03
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "context-dev/crawl-site",
"prompt": "a beautiful sunset over mountains"
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"
