bytedance/seedance/2.5/text-to-video
ByteDance's next-generation text-to-video model, generating native single-shot clips up to 30 seconds at 720p with coherent motion, native audio, and director-level camera control for professional-grade video creation.
Example output — click Run to generate your own
API README
Seedance 2.5 Text to Video
ByteDance's next-generation text-to-video model — generates native single-shot video up to 30 seconds from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.
- Need image-driven generation? Try Seedance 2.5 Image to Video.
- Need reference-guided generation? Try Seedance 2.5 Reference to Video.
Highlights
Native 30-second single-shot — Generate long, coherent takes in one request without the drift or stitching of multi-clip workflows.
Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.
Pure imagination — No input image required. Describe your scene and the model brings it to life with cinematic quality.
Flexible duration — Generate 4 to 30 second clips in a single request.
Pricing
| Resolution | Price per second |
|---|---|
| 480p | $0.2205 |
| 720p | $0.4730 |
Default: 5 seconds at 720p = $2.365. Audio generation included at no extra cost. Pricing referenced from fal.ai.
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| Creative concept visualization | Real-time video streaming |
| Social media content creation | Long-form video (>30s) |
| Marketing and ad clips | Pixel-perfect frame control |
| Storyboard-to-video workflows | Image-guided generation (use I2V) |
| Quick video prototyping | Multi-reference scenes (use Ref2V) |
Technical Specs
| Spec | Value |
|---|---|
| Input | Text prompt |
| Output | MP4 video with optional audio |
| Resolution | 480p / 720p |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 |
| Duration | 4–30 seconds |
| Execution | Async (submit → poll for result) |
| Typical latency | 2–10 minutes |
Related
- Seedance 2.5 Image to Video — Transform images into video
- Seedance 2.5 Reference to Video — Multi-reference guided generation

