Alibaba modelsvideo generation api

alibaba/wan/2.1/image-to-video

Wan 2.1 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.

Input
The text prompt to guide video generation.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the input image. If the input image does not match the chosen aspect ratio, it is resized and center cropped.
Resolution of the generated video (480p or 720p). 480p is 0.5 billing units, and 720p is 1 billing unit. Allowed values: 480p, 720p.
110
Classifier-free guidance scale. Higher values give better adherence to the prompt but may decrease quality. Range: 1 to 10.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
Random seed for reproducibility. If None, a random seed is chosen.
Idle

Example output — click Run to generate your own

API README

Wan 2.1 Image to Video

Wan 2.1 Image to Video creates a generated clip from one opening image and a motion prompt. The still supplies the scene and subject appearance, while the text explains what should happen after that first frame. It is a straightforward image-animation workflow within the Wan 2.1 family.

The route targets compact 480p or 720p outputs where the task is establishing movement from an existing visual. Guidance strength can tune how tightly generation follows the written direction. The result is delivered as a new video asset.

Highlights

  • High visual quality from images. The exact route is presented as generating high-quality video from an image. The supplied frame anchors the sequence.
  • Motion diversity. Wan 2.1 explicitly emphasizes varied motion in image-to-video results. A starting still can support different action concepts.
  • Dedicated image-to-video evaluation result. In a dedicated image-to-video human evaluation, Wan2.1 led every compared model. This statement is limited to that published comparison.
  • Natural-looking motion from a single still. Wan 2.1 produces fluid, natural-looking video and realistic motion from one input image, supporting scenes that should feel animated rather than simply displaced.

Pricing

Billing unitPrice
Base price per request$0.20

When to Use

✅ Good fit❌ Consider alternatives
Turning one established image into a moving sceneStarting without an image
Exploring movements from the same frameEditing an existing video
Producing compact 480p or 720p conceptsRequiring 1080p or 4K
Tuning adherence to a motion briefRequiring a fixed frame sequence
Creating a clip for rapid visual reviewAdding a soundtrack through this schema

Prompt Guide

Describe the motion and camera behavior that should emerge from the supplied frame.

Subject action: [movement]
Camera: [shot and motion]
Environment: [moving elements]
Pacing: [tempo]
Ending: [final framing]
{
  "image": "https://example.com/opening-image.png",
  "prompt": "The cars race forward in slow motion as the camera pans beside them and settles into a wide view.",
  "resolution": "720p",
  "guide_scale": 5,
  "aspect_ratio": "16:9",
  "seed": 42
}

Technical Specs

SpecValue
Model IDalibaba/wan/2.1/image-to-video
Required inputsimage and prompt
Resolution480p or 720p; default 720p
Guidance scale1 to 10; default 5
Aspect ratio10 presets from 21:9 through 9:16
RepeatabilityOptional integer seed
OutputVideo URL and content type
ExecutionAsynchronous job

Related

Related Models

alibaba/wan/2.1/image-to-video/loraWan 2.1 Lora is Alibaba's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.alibaba/wan/2.1/flf-to-videoWan 2.1 Flf To Video by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/text-to-videoWan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/2.1/text-to-video/loraWan 2.1 Lora by Alibaba - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.alibaba/wan/2.1/vaceWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.alibaba/wan/2.1/vace/long-reframeWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.