Bytedance modelsvideo generation api

bytedance/seedance/2.5/reference-to-video

ByteDance's next-generation reference-to-video model, generating video from multimodal references (images, videos, audio) and locking a character, set, and palette across a full take up to 30 seconds for production-grade...

Input
Idle

Example output — click Run to generate your own

API README

Seedance 2.5 Reference to Video

ByteDance's next-generation multi-reference video model — combines reference images, videos, and audio to lock a character, set, and palette across a full take up to 30 seconds for production-grade consistency.

Highlights

Multi-modal references — Combine reference images, videos, and audio in a single request. Refer to them in the prompt as @Image1, @Video1, @Audio1, etc. Up to 50 files across all modalities.

Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.

Style + identity lock — Use references to hold visual style, motion patterns, and subject identity steady across the whole clip.

Flexible duration — Generate 4 to 30 second clips in a single request.

Pricing

ResolutionPrice per second
480p$0.2205
720p$0.4730

Default: 5 seconds at 720p = $2.365. When video references are provided, both input and output video are billed (fal.ai applies a 0.6x multiplier to the per-second rate for video-input requests). Pricing referenced from fal.ai.

When to Use

✅ Good fit❌ Consider alternatives
Brand-consistent video productionReal-time video streaming
Style-guided content creationLong-form video (>30s)
Video-to-video style transferSimple text prompts (use T2V)
Audio-synced video generationSingle image animation (use I2V)
Multi-reference scene compositionQuick prototyping without assets

Input Requirements

  • Reference images/videos/audio as publicly accessible URLs
  • Supported image formats: JPG, PNG, WebP (max 30 MB, up to 30 images)
  • Supported video formats: MP4, MOV (up to 10 videos, 1.8–30.2s each, combined ≤ 30.2s)
  • Supported audio formats: MP3, WAV (up to 10 files, combined ≤ 30.2s)
  • Audio cannot be used alone — must pair with at least one reference image or video

Technical Specs

SpecValue
InputText + reference images/videos/audio
OutputMP4 video with optional audio
Resolution480p / 720p
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4, 21:9
Duration4–30 seconds
ExecutionAsync (submit → poll for result)
Typical latency2–10 minutes

Related

Related Models

bytedance/seedance/2.5/image-to-videoByteDance's next-generation image-to-video model, animating a single still into a native clip up to 30 seconds at 720p with continuous, coherent motion, native audio, and director-level camera control.bytedance/seedance/2.5/text-to-videoByteDance's next-generation text-to-video model, generating native single-shot clips up to 30 seconds at 720p with coherent motion, native audio, and director-level camera control for professional-grade video creation.bytedance/seedance/1.0/pro/image-to-videoSeedance 1.0 Pro is Bytedance's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.bytedance/seedance/1.0/pro/text-to-videoSeedance 1.0 Pro by Bytedance - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.bytedance/seedance/1.5/pro/image-to-videoByteDance Seedance v1.5 Pro image-to-video model transforming still images into cinematic video with native audio generation, camera control, and professional-grade output quality.bytedance/seedance/1.5/pro/text-to-videoByteDance Seedance v1.5 Pro text-to-video model generating cinematic video from text prompts with native audio generation, camera control, and professional-grade output quality.bytedance/seedance/2.0/fast/image-to-videoByteDance's most advanced image-to-video model in its fast tier delivering lower latency and cost without compromising on cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.bytedance/seedance/2.0/fast/reference-to-videoByteDance's most advanced reference-to-video model in its fast tier delivering lower latency and cost without compromising on cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.