API reference · NVIDIA
nvidia/cosmos-3-super/image-to-video
Integrate this model through SandBase's unified API, with production-ready schemas and examples.
Production endpoint
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
POST
https://api.sandbase.ai/v1/runModel ID
nvidia/cosmos-3-super/image-to-video01
Input Schema
12 parameters · 2 required · 10 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
image | string | Required | URL of the conditioning first-frame image for the video. |
prompt | string | Required | Text prompt describing the motion and scene of the video to generate. · Min length: 1 · Max length: 4096 |
seed | integer | Optional | The same seed and prompt given to the same model version will produce the same video every time. |
num_frames | integer | Optional | Number of frames to generate. More frames yield a longer video. · Min: 5 · Max: 189 · Default: 189 |
aspect_ratio | string | Optional | The aspect ratio of the generated image. · Options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 21:916:93:24:35:41:14:53:42:39:16 |
guidance_scale | number | Optional | Classifier-free guidance scale. Higher values increase prompt adherence at the cost of diversity. · Min: 0 · Max: 20 · Default: 6 |
frames_per_second | integer | Optional | Frames per second of the output video. · Min: 4 · Max: 60 · Default: 24 |
agentic_early_stop | boolean | Optional | Stop the agentic loop early when the critic score clears the strict quality threshold. · Default: true |
num_inference_steps | integer | Optional | Number of denoising steps. More steps yield higher quality but take longer. · Min: 1 · Max: 50 · Default: 28 |
agentic_max_iterations | integer | Optional | Maximum agentic prompt stages when agentic generation is enabled. · Min: 1 · Max: 3 · Default: 2 |
enable_agentic_generation | boolean | Optional | Enable the iterative Cosmos agentic loop: prompt upsampling, candidate video generation, VLM critique of sampled frames, and prompt rewrite. Each candidate is a full render, so this is substantially slower and costlier than a single generation. · Default: false |
agentic_samples_per_iteration | integer | Optional | Candidate videos to generate and judge per agentic iteration. The best candidate advances to the next rewrite stage. · Min: 1 · Max: 3 · Default: 2 |
02
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
03
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "nvidia/cosmos-3-super/image-to-video",
"image": "https://storage.googleapis.com/falserverless/example_inputs/hunyuan_i2v.jpg",
"prompt": "The camera slowly pushes in as the subject turns their head toward the light, hair drifting in a gentle breeze, dust motes floating through warm afternoon sun.",
"num_frames": 189,
"guidance_scale": 6,
"frames_per_second": 24,
"agentic_early_stop": true,
"num_inference_steps": 28,
"agentic_max_iterations": 2,
"enable_agentic_generation": false,
"agentic_samples_per_iteration": 2
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"
