Seedance 2.5 Reference to Video

bytedance/seedance/2.5/reference-to-video

ByteDance's next-generation reference-to-video model, generating video from multimodal references (images, videos, audio) and locking a character, set, and palette across a full take up to 30 seconds for production-grade consistency.

Base price
$1.12 / run
Execution
async
Model type
video
Input fields
8

Try the model

Playground

Open playground
Input
The text prompt to generate a video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum each

The URLs of reference images to guide video generation. Refer to them in the prompt as @Image1, @Image2, etc.
The URLs of reference videos to guide video generation. Up to 10 MP4 or MOV files; combined duration must not exceed 30 seconds.
The URLs of reference audio files to guide video generation.
The duration of the generated video in seconds (4-30). Allowed values: 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30.
The aspect ratio of the generated video. Allowed values: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9.
The resolution of the video to generate. Allowed values: 480p, 720p, 1080p.
Whether to generate synchronized audio (voice, SFX, BGM) for the video.
OutputReady

Your output will appear here

Complete the inputs, then click Run.

Specifications

Pricing

Base price
$1.12 / run
Billing formula
(usage.total_seconds ?? ((params.duration ?? 5) == 'auto' ? 5 : (params.duration ?? 5))) * ((params.resolution ?? '720p') == '480p' ? 0.099703264095 : (params.resolution ?? '720p') == '720p' ? 0.224332344214 : (params.resolution ?? '720p') == '1080p' ? 0.555222551929 : (0 / 0)) + (params.has_input_videos ? usage.input_video_seconds : 0) * ((params.resolution ?? '720p') == '480p' ? 0.059821958457 : (params.resolution ?? '720p') == '720p' ? 0.134599406528 : (params.resolution ?? '720p') == '1080p' ? 0.331691394659 : (0 / 0))

Context & modalities

Input
Schema-defined
Output
video

Capabilities

Chat
Not supported
Vision
Not supported
Reasoning
Not supported
Structured output
Not supported
Function calling
Not supported
Audio input
Not supported

Access

Provider
Bytedance
Model ID
bytedance/seedance/2.5/reference-to-video
Execution
async
API
Unified Run API
Endpoint
/v1/run

API README

Seedance 2.5 Reference to Video

ByteDance's next-generation multi-reference video model — combines reference images, videos, and audio to lock a character, set, and palette across a full take up to 30 seconds for production-grade consistency.

Highlights

Multi-modal references — Combine reference images, videos, and audio in a single request. Refer to them in the prompt as @Image1, @Video1, @Audio1, etc. Up to 50 files across all modalities.

Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.

Style + identity lock — Use references to hold visual style, motion patterns, and subject identity steady across the whole clip.

Flexible duration — Generate 4 to 30 second clips in a single request.

Pricing

Rates use Volcengine's non-promotional list price converted at 1 USD = 6.74 CNY.

ResolutionOutput price per secondInput-video price per second
480p$0.10$0.06
720p$0.22$0.13
1080p$0.56$0.33

Default: 5 output seconds at 720p = $1.12. Video references add the input-video rate using the server-probed trusted aggregate input duration. Only the exact Seedance 2.0 standard, Fast, and Mini reference-to-video modes and Seedance 2.5 reference-to-video mode, each with non-empty videos, explicitly use the isolated lenient-duration policy; the default media probe, Webhook validation, empty-video requests, and every other model or caller remain strict. Every legal domain in this isolated policy uses the system resolver, standard HTTP transport, and ProxyFromEnvironment without all-address public-DNS validation, IP pinning, or DNS-rebinding protection. Each redirect hop repeats the base URL and redirect validation. Invalid URL or host, redirect-policy, count-limit, per-item or aggregate size-limit, parser-resource-limit, and successfully parsed aggregate duration-limit failures fail closed. All other fixed failures—dns_unavailable, connect_error, network_timeout, tls_error, http_status, download_error, empty_body, mime_mismatch, magic_mismatch, unsupported_media, malformed_media, and duration_unavailable—use the 30-second family cap for the whole request once, not per video. Logs and errors expose only fixed stage/reason, video index/count, host hash, bytes, elapsed time, fallback marker, and policy version; complete URLs, queries, signatures, response bodies, and media content remain redacted. New requests freeze trusted-probe policy v4; existing v3, v2, and legacy predictions retain their frozen settlement semantics. This isolated application policy requires no new server, Kubernetes, proxy, egress, or environment configuration. Image and audio references do not add this video-input charge. Audio generation is included at no extra cost.

When to Use

✅ Good fit❌ Consider alternatives
Brand-consistent video productionReal-time video streaming
Style-guided content creationLong-form video (>30s)
Video-to-video style transferSimple text prompts (use T2V)
Audio-synced video generationSingle image animation (use I2V)
Multi-reference scene compositionQuick prototyping without assets

Input Requirements

  • Reference images/videos/audio as publicly accessible URLs
  • Supported image formats: JPG, PNG, WebP (max 30 MB, up to 30 images)
  • Supported video formats: MP4, MOV (up to 10 videos, 1.8–30.2s each, combined ≤ 30.2s)
  • Supported audio formats: MP3, WAV (up to 10 files, combined ≤ 30.2s)
  • Audio cannot be used alone — must pair with at least one reference image or video

Technical Specs

SpecValue
InputText + reference images/videos/audio
OutputMP4 video with optional audio
Resolution480p / 720p / 1080p
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4, 21:9
Duration4–30 seconds
ExecutionAsync (submit → poll for result)
Typical latency2–10 minutes

Related

Start building

Send your first request

OpenAI-compatible endpoint with unified authentication and usage tracking.

Production API
Endpoint
https://api.sandbase.ai/v1/run
Model ID
bytedance/seedance/2.5/reference-to-video
# Preserve Unicode and shell metacharacters in the request body.
result=$(printf '%s' 'ewogICJtb2RlbCI6ICJieXRlZGFuY2Uvc2VlZGFuY2UvMi41L3JlZmVyZW5jZS10by12aWRlbyIsCiAgImltYWdlcyI6IFsKICAgICJodHRwczovL3N0YXRpYy5zYW5kYmFzZS5haS9leGFtcGxlcy9ieXRlZGFuY2Uvc2VlZGFuY2UvMi4wL3JlZmVyZW5jZS10by12aWRlby9pbnB1dF9pbWFnZV9zYWZlX29jZWFuXzAuanBnIgogIF0sCiAgInByb21wdCI6ICJVc2UgSW1hZ2UxIGFzIHRoZSB2aXN1YWwgcmVmZXJlbmNlIGFuZCBWaWRlbzEgYXMgdGhlIG1vdGlvbiByZWZlcmVuY2UuIENyZWF0ZSBhIHBlYWNlZnVsIGNpbmVtYXRpYyBzaG9yZWxpbmUgc2NlbmUgd2l0aCByb2xsaW5nIG9jZWFuIHdhdmVzIGF0IGdvbGRlbiBob3VyLCBzb2Z0IGZvYW0sIG5hdHVyYWwgd2F0ZXIgbW92ZW1lbnQsIGFuZCBubyBwZW9wbGUuIiwKICAidmlkZW9zIjogWwogICAgImh0dHBzOi8vc3RhdGljLnNhbmRiYXNlLmFpL2V4YW1wbGVzL3BpeHZlcnNlL3NvdW5kLWVmZmVjdHMvaW5wdXRfdmlkZW9fMC5tcDQiCiAgXSwKICAiZHVyYXRpb24iOiA1LAogICJyZXNvbHV0aW9uIjogIjcyMHAiLAogICJnZW5lcmF0ZV9hdWRpbyI6IHRydWUKfQ==' \
  | base64 --decode \
  | curl --fail-with-body --silent \
      -X POST "https://api.sandbase.ai/v1/run" \
      -H "Authorization: Bearer $SANDBASE_API_KEY" \
      -H "Content-Type: application/json" \
      --data-binary @-)
run_id=$(printf '%s' "$result" | jq -r .id)
for attempt in $(seq 1 120); do
  status=$(printf '%s' "$result" | jq -r .status)
  case "$status" in completed|failed|timeout) break ;; esac
  sleep 2
  result=$(curl --fail-with-body --silent \
    -H "Authorization: Bearer $SANDBASE_API_KEY" \
    "https://api.sandbase.ai/v1/run/$run_id")
done
status=$(printf '%s' "$result" | jq -r .status)
[ "$status" = completed ] || { echo "Generation ended: $status" >&2; exit 1; }
printf '%s\n' "$result"

Choose your model

Compare models

All language models
ModelContextInput / 1MOutput / 1MReleased
Bytedance———Jul 1, 2026
Bytedance———Jul 1, 2026
Bytedance———Jul 1, 2026
Bytedance———Jun 16, 2025
Bytedance———Jun 16, 2025
Bytedance———Dec 23, 2025