Seedance 2.0 Text to Video

Bytedancevideo

Published April 1, 2026

About

ByteDance's most advanced text-to-video model delivering cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.

Documentation

Seedance 2.0 Text to Video

ByteD's flagship text-to-video model — generates cinematic video from text prompts with native audio, flexible duration, and director-level camera control.

Highlights

Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.

Pure imagination — No input image required. Describe your scene and the model brings it to life with cinematic quality.

Flexible duration — Generate 4 to 15 second clips in a single request. Use -1 for model-intelligent duration selection.

Up to 1080p — Full HD output with support for 6 aspect ratios including cinematic 21:9.

Pricing

ResolutionPrice per second
480p$0.10
720p$0.20
1080p$0.50

Default: 5 seconds at 720p = $1.00. Audio generation included at no extra cost.

When to Use

✅ Good fit❌ Consider alternatives
Creative concept visualizationReal-time video streaming
Social media content creationLong-form video (>15s)
Marketing and ad clipsPixel-perfect frame control
Storyboard-to-video workflowsImage-guided generation (use I2V)
Quick video prototypingMulti-reference scenes (use Ref2V)

Technical Specs

SpecValue
InputText prompt
OutputMP4 video with optional audio
Resolution480p / 720p / 1080p
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4, 21:9
Duration4–15 seconds
ExecutionAsync (submit → poll for result)
Typical latency2–8 minutes

Related

Try Seedance 2.0 Text to Video

Test this model in the Sandbase Playground with your own prompts.

Open in Playground

Related Models