alibaba/wan/2.5/text-to-video
Wan 2.5 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runalibaba/wan/2.5/text-to-videoInput Schema
6 parameters · 1 required · 5 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Required | The text prompt for video generation. Supports Chinese and English, max 800 characters. · Min length: 1 |
seed | integer | Optional | Random seed for reproducibility. If None, a random seed is chosen. |
audio | string | Optional | URL of the audio to use as the background music. Must be publicly accessible. Limit handling: If the audio duration exceeds the duration value (5 or 10 seconds), the audio is truncated to the first 5 or 10 seconds, and the rest is discarded. If the audio is shorter than the video, the remaining part of the video will be silent. For example, if the audio is 3 seconds long and the video duration is 5 seconds, the first 3 seconds of the output video will have sound, and the last 2 seconds will be silent. - Format: WAV, MP3. - Duration: 3 to 30 s. - File size: Up to 15 MB. |
duration | integer | Optional | Duration of the generated video in seconds. Choose between 5 or 10 seconds. · Options: 5, 10 · Default: 5 510 |
resolution | string | Optional | Video resolution tier · Options: 480p, 720p, 1080p · Default: "1080p" 480p720p1080p |
aspect_ratio | string | Optional | The aspect ratio of the generated image. · Options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 21:916:93:24:35:41:14:53:42:39:16 |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "alibaba/wan/2.5/text-to-video",
"prompt": "The white dragon warrior stands still, eyes full of determination and strength. The camera slowly moves closer or circles around the warrior, highlighting the powerful presence and heroic spirit of the character.",
"duration": 5,
"resolution": "1080p"
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Wan 2.5 Text-to-Video
Wan 2.5 Text-to-Video is the text-first generation route in the Wan 2.5 video family. It converts a written scene brief into a video result, giving creators a direct starting point when no existing visual asset should define the opening frame.
The endpoint supports an audio-aware workflow as well as text-only generation. This lets a creative team begin from a narrative idea or incorporate an existing soundtrack when timing and audiovisual coordination are part of the brief.
Highlights
Prompt-directed video generation. Turns written scene direction into video, using the brief to establish the subject, action, environment, and progression of the generated sequence.
Realistic motion and composition. Generates professional-quality clips with realistic motion, lighting, and scene composition, helping the result read as a coherent moving shot.
Audio-guided sequencing. Accepts an audio track as creative guidance when the generated sequence should respond to an existing soundtrack, beat, or timing reference.
Detailed prompt adherence. Follows specific creative instructions more faithfully than Wan 2.2, helping multi-part scene briefs carry through from the written direction into the generated shot.
Pricing
| Duration | 480p | 720p | 1080p |
|---|---|---|---|
| 5 seconds | $0.250 | $0.500 | $0.750 |
| 10 seconds | $0.500 | $1.000 | $1.500 |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| The scene should be created entirely from written direction | A source image must anchor the opening frame |
| A managed asynchronous result is suitable for the production pipeline | A synchronous, interactive editor is essential |
| The documented controls cover the required duration, framing, or format | The project needs controls outside this endpoint's schema |
| Creative iteration benefits from a repeatable request structure | Exact deterministic pixels, frames, geometry, or samples are mandatory |
| A finished downloadable media asset is the desired deliverable | Editable source layers or a native project file are required |
Prompt Guide
For audio transformation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.
{
"aspect_ratio": "21:9",
"audio": "https://example.com/source.wav",
"duration": 5,
"prompt": "The white dragon warrior stands still, eyes full of determination and strength. The camera slowly moves closer or circles around the warrior, highlighting the powerful presence and heroic spirit of the character.",
"resolution": "1080p"
}
Technical Specs
| Spec | Value |
|---|---|
| Model ID | alibaba/wan/2.5/text-to-video |
| Inputs | aspect_ratio, audio, duration, prompt, resolution, seed |
| Required inputs | prompt |
| Output fields | content_type, url |
| Execution | Async (submit, then poll for result) |
| Duration | 5 / 10 |
| Resolution | 480p / 720p / 1080p |
| Aspect Ratio | 21:9 / 16:9 / 3:2 / 4:3 / 5:4 / 1:1 / 4:5 / 3:4 / 2:3 / 9:16 |
Related Models
alibaba/wan/2.5/image-edit— Compare a nearby route in the same local model family.alibaba/wan/2.5/image-to-video— Compare a nearby route in the same local model family.alibaba/wan/2.5/text-to-image— Compare a nearby route in the same local model family.

