alibaba/wan/v2.2-5b/text-to-video/distill
Wan V2.2 5b Distill is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runalibaba/wan/v2.2-5b/text-to-video/distillInput Schema
10 parameters · 1 required · 9 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Required | The text prompt to guide video generation. |
seed | integer | Optional | Random seed for reproducibility. If None, a random seed is chosen. |
shift | number | Optional | Shift value for the video. Must be between 1.0 and 10.0. · Min: 1 · Max: 10 · Default: 5 |
resolution | string | Optional | Resolution of the generated video (580p or 720p). · Options: 580p, 720p · Default: "720p" 580p720p |
aspect_ratio | string | Optional | The aspect ratio of the generated image. · Options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16 21:916:93:24:35:41:14:53:42:39:16 |
guidance_scale | number | Optional | Min: 1 · Max: 10 · Default: 1 |
interpolator_model | string | Optional | The model to use for frame interpolation. If None, no interpolation is applied. · Options: none, film, rife · Default: "film" nonefilmrife |
num_inference_steps | integer | Optional | Min: 2 · Max: 50 · Default: 40 |
num_interpolated_frames | integer | Optional | Number of frames to interpolate between each pair of generated frames. Must be between 0 and 4. · Min: 0 · Max: 4 · Default: 0 |
adjust_fps_for_interpolation | boolean | Optional | If true, the number of frames per second will be multiplied by the number of interpolated frames plus one. For example, if the generated frames per second is 16 and the number of interpolated frames is 1, the final frames per second will be 32. If false, the passed frames per second will be used as-is. · Default: true |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "alibaba/wan/v2.2-5b/text-to-video/distill",
"shift": 5,
"prompt": "A medium shot establishes a modern, minimalist office setting: clean lines, muted grey walls, and polished wood surfaces. The focus shifts to a close-up on a woman in sharp, navy blue business attire. Her crisp white blouse contrasts with the deep blue of her tailored suit jacket. The subtle texture of the fabric is visible—a fine weave with a slight sheen. Her expression is serious, yet engaging, as she speaks to someone unseen just beyond the frame. Close-up on her eyes, showing the intensity of her gaze and the fine lines around them that hint at experience and focus. Her lips are slightly parted, as if mid-sentence. The light catches the subtle highlights in her auburn hair, meticulously styled. Note the slight catch of light on the silver band of her watch. High resolution 4k",
"resolution": "720p",
"guidance_scale": 1,
"interpolator_model": "film",
"num_inference_steps": 40,
"num_interpolated_frames": 0,
"adjust_fps_for_interpolation": true
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Wan 2.2 5B Distilled
Wan 2.2 5B Distilled uses the compact Wan 2.2 5B architecture in a distilled form, reducing the work needed to produce prompt-led video while retaining the family’s scene, motion, and composition knowledge. It is aimed at efficient generation rather than maximizing model size or inference complexity.
Write a concise cinematic brief that establishes the central action, framing, lighting, and atmosphere without burying the model in secondary detail. This version suits rapid experiments, resource-conscious pipelines, teaching environments, and batch concept generation where a smaller distilled system is preferable.
Highlights
- Distilled video generation. Compresses learned video behavior into a more efficient inference process.
- Compact 5B architecture. Reduces model scale while retaining prompt-to-motion capability.
- Prompt-led scene construction. Builds subjects, environment, action, and camera from written direction.
- Efficient concept iteration. Supports higher-throughput exploration and resource-conscious deployment.
Pricing
| Billing unit | Price |
|---|---|
| Per request | $0.08 |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| The scene should be created entirely from written direction | A source image must anchor the opening frame |
| A managed asynchronous result is suitable for the production pipeline | A synchronous, interactive editor is essential |
| The documented controls cover the required duration, framing, or format | The project needs controls outside this endpoint's schema |
| Creative iteration benefits from a repeatable request structure | Exact deterministic pixels, frames, geometry, or samples are mandatory |
| A finished downloadable media asset is the desired deliverable | Editable source layers or a native project file are required |
Prompt Guide
For generation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.
{
"aspect_ratio": "21:9",
"prompt": "A medium shot establishes a modern, minimalist office setting: clean lines, muted grey walls, and polished wood surfaces. The focus shifts to a close-up on a woman in sharp, navy blue business attire. Her crisp white blouse contrasts with the deep blue of her tailored suit jacket. The subtle texture of the fabric is visible—a fine weave with a slight sheen. Her expression is serious, yet engaging, as she speaks to someone unseen just beyond the frame. Close-up on her eyes, showing the intensity of her gaze and the fine lines around them that hint at experience and focus. Her lips are slightly parted, as if mid-sentence. The light catches the subtle highlights in her auburn hair, meticulously styled. Note the slight catch of light on the silver band of her watch. High resolution 4k",
"resolution": "720p"
}
Technical Specs
| Spec | Value |
|---|---|
| Model ID | alibaba/wan/v2.2-5b/text-to-video/distill |
| Inputs | adjust_fps_for_interpolation, aspect_ratio, guidance_scale, interpolator_model, num_inference_steps, num_interpolated_frames, prompt, resolution, seed, shift |
| Required inputs | prompt |
| Output fields | content_type, url |
| Execution | Async (submit, then poll for result) |
| Resolution | 580p / 720p |
| Aspect Ratio | 21:9 / 16:9 / 3:2 / 4:3 / 5:4 / 1:1 / 4:5 / 3:4 / 2:3 / 9:16 |
Related Models
alibaba/wan/2.1/flf-to-video— Compare this concrete local family route.alibaba/wan/2.1/image-to-video— Compare this concrete local family route.alibaba/wan/2.1/image-to-video/lora— Compare this concrete local family route.

