SandBase is live — $1 in free credits on signupStart free ›

Bytedance modelsvideo generation api

bytedance/seedance/2.0/reference-to-video

ByteDance's most advanced reference-to-video model generating cinematic video guided by reference content, with native audio, multi-shot editing, and director-level camera control for professional-grade video creation.

Input
The text prompt to generate a video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum each

The URLs of reference images to guide video generation.
The URLs of reference videos to guide video generation. Up to 3 MP4 or MOV files, 2-15 seconds combined duration, under 50 MB total.
The URLs of reference audio files to guide video generation.
The duration of the generated video in seconds. Allowed values: 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
The aspect ratio of the generated video. Allowed values: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9.
The resolution of the video to generate. Allowed values: 480p, 720p, 1080p, 4k.
Idle

Example output — click Run to generate your own

API README

Seedance 2.0 Reference to Video

ByteD's flagship multi-reference video generation model — combines images, videos, and audio references to produce cinematic output with precise style and motion control.

Highlights

Multi-modal references — Combine reference images, videos, and audio in a single request for rich, guided generation.

Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.

Style transfer — Use reference images and videos to guide the visual style, motion patterns, and composition.

Flexible duration — Generate 4 to 15 second clips in a single request. Use -1 for model-intelligent duration selection.

Up to 4K — 4K output with support for 6 aspect ratios including cinematic 21:9.

Pricing

Non-4K rates use Volcengine's non-promotional list price converted at 1 USD = 6.74 CNY. All 4K output rates use the confirmed reference price.

ResolutionOutput price per secondInput-video price per second
480p$0.07$0.04
720p$0.15$0.09
1080p$0.37$0.22
4K$1.80$1.11

Default: 5 output seconds at 720p = $0.74. Video references add the input-video rate using the server-probed trusted aggregate input duration. Only the exact Seedance 2.0 standard, Fast, and Mini reference-to-video modes and Seedance 2.5 reference-to-video mode, each with non-empty videos, explicitly use the isolated lenient-duration policy; the default media probe, Webhook validation, empty-video requests, and every other model or caller remain strict. Every legal domain in this isolated policy uses the system resolver, standard HTTP transport, and ProxyFromEnvironment without all-address public-DNS validation, IP pinning, or DNS-rebinding protection. Each redirect hop repeats the base URL and redirect validation. Invalid URL or host, redirect-policy, count-limit, per-item or aggregate size-limit, parser-resource-limit, and successfully parsed aggregate duration-limit failures fail closed. All other fixed failures—dns_unavailable, connect_error, network_timeout, tls_error, http_status, download_error, empty_body, mime_mismatch, magic_mismatch, unsupported_media, malformed_media, and duration_unavailable—use the 15-second family cap for the whole request once, not per video. Logs and errors expose only fixed stage/reason, video index/count, host hash, bytes, elapsed time, fallback marker, and policy version; complete URLs, queries, signatures, response bodies, and media content remain redacted. New requests freeze trusted-probe policy v4; existing v3, v2, and legacy predictions retain their frozen settlement semantics. This isolated application policy requires no new server, Kubernetes, proxy, egress, or environment configuration. Image and audio references do not add this video-input charge. Audio generation is included at no extra cost.

When to Use

✅ Good fit❌ Consider alternatives
Brand-consistent video productionReal-time video streaming
Style-guided content creationLong-form video (>15s)
Video-to-video style transferSimple text prompts (use T2V)
Audio-synced video generationSingle image animation (use I2V)
Multi-reference scene compositionQuick prototyping without assets

Input Requirements

  • Reference images/videos/audio as publicly accessible URLs
  • Supported image formats: PNG, JPEG, WebP (< 5MB)
  • Supported video formats: MP4 (< 50MB)
  • Supported audio formats: MP3, WAV (< 50MB)
  • Audio cannot be used alone — must pair with image, video, or text

Technical Specs

SpecValue
InputText + reference images/videos/audio
OutputMP4 video with optional audio
Resolution480p / 720p / 1080p / 4K
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4, 21:9
Duration4–15 seconds
ExecutionAsync (submit → poll for result)
Typical latency2–8 minutes

Related

Related Models

bytedance/seedance/2.0/fast/image-to-videoByteDance's most advanced image-to-video model in its fast tier delivering lower latency and cost without compromising on cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.bytedance/seedance/2.0/fast/reference-to-videoByteDance's most advanced reference-to-video model in its fast tier delivering lower latency and cost without compromising on cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.bytedance/seedance/2.0/fast/text-to-videoByteDance's most advanced text-to-video model in its fast tier delivering lower latency and cost without compromising on cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.bytedance/seedance/2.0/image-to-videoByteDance's most advanced image-to-video model transforming still images into cinematic video with native audio, multi-shot editing, and director-level camera control for professional-grade video creation.bytedance/seedance/2.0/mini/image-to-videoSeedance 2.0 Mini by Bytedance - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.bytedance/seedance/2.0/mini/reference-to-videoSeedance 2.0 Mini by Bytedance - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.bytedance/seedance/2.0/mini/text-to-videoSeedance 2.0 Mini is Bytedance's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.bytedance/seedance/2.0/text-to-videoByteDance's most advanced text-to-video model delivering cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.