SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/v2.2-5b/text-to-video/distill

Wan V2.2 5b Distill is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.

Input
The text prompt to guide video generation.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
Resolution of the generated video (580p or 720p). Allowed values: 580p, 720p.
Random seed for reproducibility. If None, a random seed is chosen.
The model to use for frame interpolation. If None, no interpolation is applied. Allowed values: none, film, rife.
04
Number of frames to interpolate between each pair of generated frames. Must be between 0 and 4. Range: 0 to 4.
If true, the number of frames per second will be multiplied by the number of interpolated frames plus one. For example, if the generated frames per second is 16 and the number of interpolated frames is 1, the final frames per second will be 32. If false, the passed frames per second will be used as-is.
110
Shift value for the video. Must be between 1.0 and 10.0. Range: 1 to 10.
110
Range: 1 to 10.
250
Range: 2 to 50.
Idle

Example output — click Run to generate your own

API README

Wan 2.2 5B Distilled

Wan 2.2 5B Distilled uses the compact Wan 2.2 5B architecture in a distilled form, reducing the work needed to produce prompt-led video while retaining the family’s scene, motion, and composition knowledge. It is aimed at efficient generation rather than maximizing model size or inference complexity.

Write a concise cinematic brief that establishes the central action, framing, lighting, and atmosphere without burying the model in secondary detail. This version suits rapid experiments, resource-conscious pipelines, teaching environments, and batch concept generation where a smaller distilled system is preferable.

Highlights

  • Distilled video generation. Compresses learned video behavior into a more efficient inference process.
  • Compact 5B architecture. Reduces model scale while retaining prompt-to-motion capability.
  • Prompt-led scene construction. Builds subjects, environment, action, and camera from written direction.
  • Efficient concept iteration. Supports higher-throughput exploration and resource-conscious deployment.

Pricing

Billing unitPrice
Per request$0.08

When to Use

✅ Good fit❌ Consider alternatives
The scene should be created entirely from written directionA source image must anchor the opening frame
A managed asynchronous result is suitable for the production pipelineA synchronous, interactive editor is essential
The documented controls cover the required duration, framing, or formatThe project needs controls outside this endpoint's schema
Creative iteration benefits from a repeatable request structureExact deterministic pixels, frames, geometry, or samples are mandatory
A finished downloadable media asset is the desired deliverableEditable source layers or a native project file are required

Prompt Guide

For generation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.

{
  "aspect_ratio": "21:9",
  "prompt": "A medium shot establishes a modern, minimalist office setting: clean lines, muted grey walls, and polished wood surfaces. The focus shifts to a close-up on a woman in sharp, navy blue business attire. Her crisp white blouse contrasts with the deep blue of her tailored suit jacket. The subtle texture of the fabric is visible—a fine weave with a slight sheen. Her expression is serious, yet engaging, as she speaks to someone unseen just beyond the frame. Close-up on her eyes, showing the intensity of her gaze and the fine lines around them that hint at experience and focus. Her lips are slightly parted, as if mid-sentence. The light catches the subtle highlights in her auburn hair, meticulously styled. Note the slight catch of light on the silver band of her watch. High resolution 4k",
  "resolution": "720p"
}

Technical Specs

SpecValue
Model IDalibaba/wan/v2.2-5b/text-to-video/distill
Inputsadjust_fps_for_interpolation, aspect_ratio, guidance_scale, interpolator_model, num_inference_steps, num_interpolated_frames, prompt, resolution, seed, shift
Required inputsprompt
Output fieldscontent_type, url
ExecutionAsync (submit, then poll for result)
Resolution580p / 720p
Aspect Ratio21:9 / 16:9 / 3:2 / 4:3 / 5:4 / 1:1 / 4:5 / 3:4 / 2:3 / 9:16

Related Models

Related Models

alibaba/wan/2.1/flf-to-videoWan 2.1 Flf To Video by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-videoWan 2.1 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-video/loraWan 2.1 Lora is Alibaba's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.alibaba/wan/2.1/text-to-videoWan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/2.1/text-to-video/loraWan 2.1 Lora by Alibaba - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.alibaba/wan/2.1/vaceWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.