SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/2.5/text-to-video

Wan 2.5 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.

Input
The text prompt for video generation. Supports Chinese and English, max 800 characters.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
URL of the audio to use as the background music. Must be publicly accessible. Limit handling: If the audio duration exceeds the duration value (5 or 10 seconds), the audio is truncated to the first 5 or 10 seconds, and the rest is discarded. If the audio is shorter than the video, the remaining part of the video will be silent. For example, if the audio is 3 seconds long and the video duration is 5 seconds, the first 3 seconds of the output video will have sound, and the last 2 seconds will be silent. - Format: WAV, MP3. - Duration: 3 to 30 s. - File size: Up to 15 MB.
Video resolution tier Allowed values: 480p, 720p, 1080p.
Duration of the generated video in seconds. Choose between 5 or 10 seconds. Allowed values: 5, 10.
Random seed for reproducibility. If None, a random seed is chosen.
Idle

Example output — click Run to generate your own

API README

Wan 2.5 Text-to-Video

Wan 2.5 Text-to-Video is the text-first generation route in the Wan 2.5 video family. It converts a written scene brief into a video result, giving creators a direct starting point when no existing visual asset should define the opening frame.

The endpoint supports an audio-aware workflow as well as text-only generation. This lets a creative team begin from a narrative idea or incorporate an existing soundtrack when timing and audiovisual coordination are part of the brief.

Highlights

Prompt-directed video generation. Turns written scene direction into video, using the brief to establish the subject, action, environment, and progression of the generated sequence.

Realistic motion and composition. Generates professional-quality clips with realistic motion, lighting, and scene composition, helping the result read as a coherent moving shot.

Audio-guided sequencing. Accepts an audio track as creative guidance when the generated sequence should respond to an existing soundtrack, beat, or timing reference.

Detailed prompt adherence. Follows specific creative instructions more faithfully than Wan 2.2, helping multi-part scene briefs carry through from the written direction into the generated shot.

Pricing

Duration480p720p1080p
5 seconds$0.250$0.500$0.750
10 seconds$0.500$1.000$1.500

When to Use

✅ Good fit❌ Consider alternatives
The scene should be created entirely from written directionA source image must anchor the opening frame
A managed asynchronous result is suitable for the production pipelineA synchronous, interactive editor is essential
The documented controls cover the required duration, framing, or formatThe project needs controls outside this endpoint's schema
Creative iteration benefits from a repeatable request structureExact deterministic pixels, frames, geometry, or samples are mandatory
A finished downloadable media asset is the desired deliverableEditable source layers or a native project file are required

Prompt Guide

For audio transformation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.

{
  "aspect_ratio": "21:9",
  "audio": "https://example.com/source.wav",
  "duration": 5,
  "prompt": "The white dragon warrior stands still, eyes full of determination and strength. The camera slowly moves closer or circles around the warrior, highlighting the powerful presence and heroic spirit of the character.",
  "resolution": "1080p"
}

Technical Specs

SpecValue
Model IDalibaba/wan/2.5/text-to-video
Inputsaspect_ratio, audio, duration, prompt, resolution, seed
Required inputsprompt
Output fieldscontent_type, url
ExecutionAsync (submit, then poll for result)
Duration5 / 10
Resolution480p / 720p / 1080p
Aspect Ratio21:9 / 16:9 / 3:2 / 4:3 / 5:4 / 1:1 / 4:5 / 3:4 / 2:3 / 9:16

Related Models

Related Models

alibaba/wan/2.5/image-to-videoWan 2.5 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-videoWan 2.1 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-video/loraWan 2.1 Lora is Alibaba's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.alibaba/wan/2.1/text-to-videoWan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/2.1/text-to-video/loraWan 2.1 Lora by Alibaba - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.alibaba/wan/2.1/vaceWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.