SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/2.1/text-to-video

Wan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.

Input
The text prompt to guide video generation.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
Resolution of the generated video (480p, 580p, or 720p). Allowed values: 480p, 720p.
Random seed for reproducibility. If None, a random seed is chosen.
If true, the video will be generated faster with no noticeable degradation in the visual quality.
Idle

Example output — click Run to generate your own

API README

Wan-2.1 Text-to-Video

Wan-2.1 Text-to-Video is a video-generation route in Alibaba’s Wan family, built on a video-focused diffusion-transformer system and a causal video VAE. It creates a complete moving scene from a required text prompt.

Wan is suited to creative briefs that need coherent temporal development rather than a disconnected set of frames. Describe subjects, actions, environment, camera language, lighting, pacing, and visual style together; the local route then exposes its supported resolution, framing, speed, and reproducibility options.

Highlights

  • Coherent video synthesis. Uses a video-focused architecture to maintain scene and motion information over time.
  • Chinese and English visual text. The Wan 2.1 family explicitly supports generation of visible text in both languages.
  • Temporal representation. Its causal video VAE is designed to preserve temporal information across generated frames.
  • Prompt-led filmmaking. Builds subject, action, setting, and camera direction from a written scene description.

Pricing

ConfigurationBilling unitPrice
Base generationPer request$0.2

When to Use

✅ Good fit❌ Consider alternatives
The project needs creating a directed video sequenceThe goal is a different media task or endpoint
The available inputs match the required local schemaRequired source media or permissions are unavailable
The brief can specify subject, composition, style, and deliveryThe result must be deterministic at pixel or sample level
The supported formats and controls match final placementDelivery requires unsupported dimensions, codecs, or duration
An asynchronous generated result fits the workflowA live, frame-synchronous, or real-time response is mandatory

Prompt Guide

Write the request as a production brief: identify the main subject or source material, state the intended transformation, describe composition or timing, and finish with style, atmosphere, and delivery constraints. Keep preservation requirements separate from requested changes, and use only fields exposed by this route.

{
  "prompt": "A cinematic, precisely directed scene with clear subject action, camera movement, lighting, and atmosphere",
  "seed": 1,
  "resolution": "480p",
  "turbo_mode": false
}

Technical Specs

SpecValue
Model IDalibaba/wan/2.1/text-to-video
Input fieldsseed (integer)<br>prompt (string)<br>resolution (string; 480p, 720p)<br>turbo_mode (boolean)<br>aspect_ratio (string; 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16)
Required inputprompt
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

  • alibaba/qwen-image-3/edit
  • alibaba/happy-horse/video-edit
  • alibaba/wan/2.2/image-to-video/turbo

Related Models

alibaba/wan/2.1/text-to-video/loraWan 2.1 Lora by Alibaba - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.alibaba/wan/2.1/image-to-videoWan 2.1 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-video/loraWan 2.1 Lora is Alibaba's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.alibaba/wan/2.1/flf-to-videoWan 2.1 Flf To Video by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/vaceWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.alibaba/wan/2.1/vace/long-reframeWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.