SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/alpha

Wan Alpha is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.

Input
The prompt to guide the video generation.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
The resolution of the generated video. Allowed values: 240p, 360p, 480p, 720p.
The seed for the random number generator.
The output type of the generated video. Allowed values: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif).
Whether to binarize the mask.
160
The frame rate of the generated video. Range: 1 to 60.
01
The threshold for mask binarization. When binarize_mask is True, this threshold will be used to binarize the mask. This will also be used for transparency when the output type is `.webm`. Range: 0 to 1.
01
The upper bound of the mask clamping. Range: 0 to 1.
01
The lower bound of the mask clamping. Range: 0 to 1.
Idle

Example output — click Run to generate your own

API README

Wan Alpha

Wan Alpha creates professional text-to-video clips from scripts and descriptive prompts, emphasizing realistic motion, lighting, and scene composition. It serves as an exploratory Wan generation model for converting narrative ideas into moving visual concepts without requiring a source image or reference clip.

Use a complete scene description with subject, action, environment, camera position, movement, time of day, lighting, and mood. It works best for concept development and cinematic ideation where the creator wants the model to invent the full visual world rather than preserve an existing asset.

Highlights

  • Script-to-scene translation. Turns written narrative direction into a staged moving sequence.
  • Realistic motion rendering. Produces subject and environmental movement intended to feel physically connected.
  • Lighting-aware cinematography. Builds illumination and atmosphere into the generated scene rather than adding them as flat style labels.
  • Composed visual storytelling. Coordinates subject placement, background, camera, and action into a readable shot.

Pricing

ResolutionPrice per generated second
240p$0.010
360p$0.010
480p$0.020
720p$0.040

When to Use

✅ Good fit❌ Consider alternatives
The project needs this exact multimodal creation workflowThe intended task belongs to a different media route
Available source media matches every required fieldRequired assets or usage rights are unavailable
The brief can define subject, action, camera, and styleOutput must be deterministic at frame or pixel level
Supported duration, resolution, and framing fit deliveryFinal placement requires unsupported specifications
An asynchronous generation job fits productionA live or frame-synchronous response is mandatory

Prompt Guide

Describe the result as a shot or design brief: identify subjects and references, state the action or transformation, specify environment and composition, then add camera behavior, lighting, pacing, style, sound, and preservation constraints where relevant. Use exact reference identifiers exposed by the local schema.

{
  "fps": 16,
  "prompt": "A cinematic scene with clearly directed subject action, camera movement, lighting, pacing, and atmosphere",
  "resolution": "240p",
  "seed": 1
}

Technical Specs

SpecValue
Model IDalibaba/wan/alpha
Input fieldsfps (integer; 1–60)<br>seed (integer)<br>prompt (string)<br>resolution (string; 240p, 360p, 480p, 720p)<br>aspect_ratio (string; 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16)<br>binarize_mask (boolean)<br>mask_clamp_lower (number; 0–1)<br>mask_clamp_upper (number; 0–1)<br>video_output_type (string; X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif))<br>mask_binarization_threshold (number; 0–1)
Required inputprompt
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

Related Models

alibaba/wan/2.1/flf-to-videoWan 2.1 Flf To Video by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-videoWan 2.1 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-video/loraWan 2.1 Lora is Alibaba's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.alibaba/wan/2.1/text-to-videoWan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/2.1/text-to-video/loraWan 2.1 Lora by Alibaba - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.alibaba/wan/2.1/vaceWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.