bytedance/seedance/2.5/reference-to-video
ByteDance's next-generation reference-to-video model, generating video from multimodal references (images, videos, audio) and locking a character, set, and palette across a full take up to 30 seconds for production-grade...
Example output — click Run to generate your own
API README
Seedance 2.5 Reference to Video
ByteDance's next-generation multi-reference video model — combines reference images, videos, and audio to lock a character, set, and palette across a full take up to 30 seconds for production-grade consistency.
- Need text-only generation? Try Seedance 2.5 Text to Video.
- Need single-image driven generation? Try Seedance 2.5 Image to Video.
Highlights
Multi-modal references — Combine reference images, videos, and audio in a single request. Refer to them in the prompt as @Image1, @Video1, @Audio1, etc. Up to 50 files across all modalities.
Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.
Style + identity lock — Use references to hold visual style, motion patterns, and subject identity steady across the whole clip.
Flexible duration — Generate 4 to 30 second clips in a single request.
Pricing
| Resolution | Price per second |
|---|---|
| 480p | $0.2205 |
| 720p | $0.4730 |
Default: 5 seconds at 720p = $2.365. When video references are provided, both input and output video are billed (fal.ai applies a 0.6x multiplier to the per-second rate for video-input requests). Pricing referenced from fal.ai.
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| Brand-consistent video production | Real-time video streaming |
| Style-guided content creation | Long-form video (>30s) |
| Video-to-video style transfer | Simple text prompts (use T2V) |
| Audio-synced video generation | Single image animation (use I2V) |
| Multi-reference scene composition | Quick prototyping without assets |
Input Requirements
- Reference images/videos/audio as publicly accessible URLs
- Supported image formats: JPG, PNG, WebP (max 30 MB, up to 30 images)
- Supported video formats: MP4, MOV (up to 10 videos, 1.8–30.2s each, combined ≤ 30.2s)
- Supported audio formats: MP3, WAV (up to 10 files, combined ≤ 30.2s)
- Audio cannot be used alone — must pair with at least one reference image or video
Technical Specs
| Spec | Value |
|---|---|
| Input | Text + reference images/videos/audio |
| Output | MP4 video with optional audio |
| Resolution | 480p / 720p |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 |
| Duration | 4–30 seconds |
| Execution | Async (submit → poll for result) |
| Typical latency | 2–10 minutes |
Related
- Seedance 2.5 Text to Video — Generate video from text prompts
- Seedance 2.5 Image to Video — Single image to video

