SandBase is live — $1 in free credits on signupStart free ›
Use in agentVidu models

Vidu modelsvideo generation api

vidu/q3/image-to-video

Vidu Q3 by Shengshu - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.

Input
Text prompt for video generation, max 2000 characters

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL or base64 image to use as the starting frame

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the image to use as the ending frame. When provided, generates a transition video between start and end frames.
Whether to use direct audio-video generation. When true, outputs video with sound (including dialogue and sound effects).
Output video resolution. Note: 360p is not available when end_image_url is provided. Allowed values: 360p, 540p, 720p, 1080p.
116
Duration of the video in seconds (1-16 for Q3 models) Range: 1 to 16.
Random seed for reproducibility. If None, a random seed is chosen.
Idle

Example output — click Run to generate your own

API README

Vidu Q3 Image to Video

Vidu Q3 Image to Video is a specialized SandBase endpoint built to turn a source still into a directed moving shot while preserving its visual anchor. It is most useful when teams need start-and-end-frame transitions, native sound generation, and flexible 1–16 second timing, with the route’s inputs defining a repeatable production contract instead of leaving critical delivery choices to an ad-hoc manual workflow.

Use this exact route when its input mode matches the creative asset already in hand: describe the visual or spoken result clearly, set only the controls that support the intended delivery, and keep the subject, action, environment, camera or performance direction internally consistent. The result is returned asynchronously as a downloadable media URL suitable for review, automation, or downstream finishing.

Highlights

  • Still-led motion. The source image establishes composition and subject appearance while the prompt directs action, camera travel, and atmosphere.
  • Optional destination frame. Add an ending image to shape a deliberate visual transition instead of leaving the final composition entirely open.
  • Synchronized audiovisual output. Native audio mode can generate dialogue, ambience, and effects together with the moving picture.
  • Production-length control. Durations from 1 to 16 seconds and four resolution tiers support everything from quick motion tests to polished clips.

Pricing

Billing follows the selected duration and resolution. The standard rate is $0.0700 per second at 360p/540p and $0.1540 per second at 720p/1080p.

DurationResolutionPrice
1 sec360p$0.0700
1 sec540p$0.0700
1 sec720p$0.1540
1 sec1080p$0.1540
2 sec360p$0.1400
2 sec540p$0.1400
2 sec720p$0.3080
2 sec1080p$0.3080
3 sec360p$0.2100
3 sec540p$0.2100
3 sec720p$0.4620
3 sec1080p$0.4620
4 sec360p$0.2800
4 sec540p$0.2800
4 sec720p$0.6160
4 sec1080p$0.6160
5 sec360p$0.3500
5 sec540p$0.3500
5 sec720p$0.7700
5 sec1080p$0.7700
6 sec360p$0.4200
6 sec540p$0.4200
6 sec720p$0.9240
6 sec1080p$0.9240
7 sec360p$0.4900
7 sec540p$0.4900
7 sec720p$1.0780
7 sec1080p$1.0780
8 sec360p$0.5600
8 sec540p$0.5600
8 sec720p$1.2320
8 sec1080p$1.2320
9 sec360p$0.6300
9 sec540p$0.6300
9 sec720p$1.3860
9 sec1080p$1.3860
10 sec360p$0.7000
10 sec540p$0.7000
10 sec720p$1.5400
10 sec1080p$1.5400
11 sec360p$0.7700
11 sec540p$0.7700
11 sec720p$1.6940
11 sec1080p$1.6940
12 sec360p$0.8400
12 sec540p$0.8400
12 sec720p$1.8480
12 sec1080p$1.8480
13 sec360p$0.9100
13 sec540p$0.9100
13 sec720p$2.0020
13 sec1080p$2.0020
14 sec360p$0.9800
14 sec540p$0.9800
14 sec720p$2.1560
14 sec1080p$2.1560
15 sec360p$1.0500
15 sec540p$1.0500
15 sec720p$2.3100
15 sec1080p$2.3100
16 sec360p$1.1200
16 sec540p$1.1200
16 sec720p$2.4640
16 sec1080p$2.4640

When to Use

ScenarioWhy this route fits
Concept developmentChoose Vidu Q3 Image to Video when its image to video workflow matches the starting material and you need several clearly directed variations.
Production iterationUse explicit duration, resolution, ratio, seed, or quality controls to compare versions without changing the core creative brief.
Channel adaptationGenerate directly in the landscape, square, portrait, or vertical format required by the destination whenever that control is available.
Automated pipelinesIntegrate the asynchronous media URL into review queues, asset libraries, publishing tools, or a later finishing stage.
Alternative routePick a related text-, image-, reference-, edit-, or turbo route when the available source media or required degree of control is different.

Prompt Guide

Lead with the main subject or source asset, then describe the intended action or transformation, environment, composition, camera or vocal delivery, lighting and mood. Keep instructions concrete and compatible; use the route’s explicit fields for duration, resolution, ratio, quality, voice, or reproducibility instead of burying those settings in prose.

{
  "prompt": "A cinematic product reveal with deliberate subject motion, coherent lighting, and a slow camera push.",
  "image": "https://example.com/input.jpg",
  "duration": 5,
  "resolution": "720p",
  "audio": true
}

Technical Specs

PropertyDetails
Model IDvidu/q3/image-to-video
Required inputsprompt, image
ExecutionAsynchronous; poll the returned generation ID
OutputDownloadable media URL
seedinteger; optional
audioboolean; optional; default: true
imagestring; required
promptstring; required; default:
durationinteger; optional; range: 1–16; default: 5
end_imagestring; optional
resolutionstring; optional; choices: 360p, 540p, 720p, 1080p; default: 720p

Related Models

ModelBest for
vidu/q3/text-to-videoAlternative text to video workflow
vidu/q3/reference-to-video/mixAlternative mix workflow
vidu/q3/image-to-video/turboAlternative turbo workflow
vidu/q3/text-to-video/turboAlternative turbo workflow

Related Models

vidu/q3/image-to-video/turboVidu Q3 Turbo by Shengshu - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.vidu/q3/reference-to-video/mixVidu Q3 Mix by Shengshu - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.vidu/q3/text-to-videoVidu Q3 is Shengshu's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.vidu/q3/text-to-video/turboVidu Q3 Turbo is Shengshu's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.vidu/q1/image-to-videoVidu Q1 by Shengshu - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.vidu/q1/reference-to-videoVidu Q1 by Shengshu - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.vidu/q1/start-end-to-videoVidu Q1 Start End To Video by Shengshu - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.vidu/q1/text-to-videoVidu Q1 is Shengshu's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.