alibaba/wan/2.1/text-to-video
Wan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runalibaba/wan/2.1/text-to-videoInput Schema
5 parameters · 1 required · 4 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Required | The text prompt to guide video generation. |
seed | integer | Optional | Random seed for reproducibility. If None, a random seed is chosen. |
resolution | string | Optional | Resolution of the generated video (480p, 580p, or 720p). · Options: 480p, 720p · Default: "720p" 480p720p |
turbo_mode | boolean | Optional | If true, the video will be generated faster with no noticeable degradation in the visual quality. · Default: false |
aspect_ratio | string | Optional | The aspect ratio of the generated image. · Options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 21:916:93:24:35:41:14:53:42:39:16 |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "alibaba/wan/2.1/text-to-video",
"prompt": "A stylish woman walks down a Tokyo street filled with warm glowing neon and animated city signage. She wears a black leather jacket, a long red dress, and black boots, and carries a black purse.",
"resolution": "720p",
"turbo_mode": false
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Wan-2.1 Text-to-Video
Wan-2.1 Text-to-Video is a video-generation route in Alibaba’s Wan family, built on a video-focused diffusion-transformer system and a causal video VAE. It creates a complete moving scene from a required text prompt.
Wan is suited to creative briefs that need coherent temporal development rather than a disconnected set of frames. Describe subjects, actions, environment, camera language, lighting, pacing, and visual style together; the local route then exposes its supported resolution, framing, speed, and reproducibility options.
Highlights
- Coherent video synthesis. Uses a video-focused architecture to maintain scene and motion information over time.
- Chinese and English visual text. The Wan 2.1 family explicitly supports generation of visible text in both languages.
- Temporal representation. Its causal video VAE is designed to preserve temporal information across generated frames.
- Prompt-led filmmaking. Builds subject, action, setting, and camera direction from a written scene description.
Pricing
| Configuration | Billing unit | Price |
|---|---|---|
| Base generation | Per request | $0.2 |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| The project needs creating a directed video sequence | The goal is a different media task or endpoint |
| The available inputs match the required local schema | Required source media or permissions are unavailable |
| The brief can specify subject, composition, style, and delivery | The result must be deterministic at pixel or sample level |
| The supported formats and controls match final placement | Delivery requires unsupported dimensions, codecs, or duration |
| An asynchronous generated result fits the workflow | A live, frame-synchronous, or real-time response is mandatory |
Prompt Guide
Write the request as a production brief: identify the main subject or source material, state the intended transformation, describe composition or timing, and finish with style, atmosphere, and delivery constraints. Keep preservation requirements separate from requested changes, and use only fields exposed by this route.
{
"prompt": "A cinematic, precisely directed scene with clear subject action, camera movement, lighting, and atmosphere",
"seed": 1,
"resolution": "480p",
"turbo_mode": false
}
Technical Specs
| Spec | Value |
|---|---|
| Model ID | alibaba/wan/2.1/text-to-video |
| Input fields | seed (integer)<br>prompt (string)<br>resolution (string; 480p, 720p)<br>turbo_mode (boolean)<br>aspect_ratio (string; 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16) |
| Required input | prompt |
| Output fields | url, content_type |
| Execution | Asynchronous job |
Related Models
alibaba/qwen-image-3/editalibaba/happy-horse/video-editalibaba/wan/2.2/image-to-video/turbo

