Alibaba modelsvideo generation api

alibaba/wan/3.0/image-to-video

Wan 3.0 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.

Input
Text prompt describing the motion to generate.
First frame of the generated video.
Last frame of the generated video. Requires start_image_url.
Include generated audio.
Output video resolution tier. Allowed values: 480p, 720p, 1080p.
230
Output duration in seconds. Set to null for smart duration, which lets the model pick a length from the prompt and reference media. Range: 2 to 30.
Enable enhanced reasoning before generation.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
02147483647
Range: 0 to 2147483647.
Idle

Example output — click Run to generate your own

API README

Wan 3.0

Wan 3.0 uses a supplied opening image to produce a new sequence with the latest 3.0 family’s emphasis on coherent motion, scene detail, and cinematic presentation. Its creative focus is preserving the source composition while introducing controlled movement, while balancing generation quality with the standard 3.0 cost profile.

Plan the clip around a clear temporal arc: define the initial state, meaningful action, camera behavior, environmental response, and intended ending. Avoid redescribing the still; concentrate on what begins moving and how the camera reveals it.

Highlights

  • 3.0 image-conditioned motion. Develops a coherent moving shot from the visual identity and layout of the opening still.
  • Improved motion coherence. Maintains more stable subject action and environmental dynamics across the generated clip.
  • Detailed standard rendering. Balances spatial detail and cinematic presentation within the standard 3.0 model.
  • Reference-composition continuity. Preserves important source geometry and identity while the shot evolves.

Pricing

ConfigurationBilling unitPrice
Generated duration480p, per second$0.050
Generated duration720p, per second$0.100
Generated duration1080p, per second$0.200

When to Use

✅ Good fit❌ Consider alternatives
The project needs this exact image-to-video workflowThe intended task belongs to a different media route
Available source media matches every required fieldRequired assets or usage rights are unavailable
The brief can define subject, action, camera, and styleOutput must be deterministic at frame or pixel level
Supported duration, resolution, and framing fit deliveryFinal placement requires unsupported specifications
An asynchronous generation job fits productionA live or frame-synchronous response is mandatory

Prompt Guide

Describe the result as a shot or design brief: identify subjects and references, state the action or transformation, specify environment and composition, then add camera behavior, lighting, pacing, style, sound, and preservation constraints where relevant. Use exact reference identifiers exposed by the local schema.

{
  "audio": true,
  "end_image": "https://example.com/reference.jpg",
  "image": "https://example.com/start-frame.png",
  "prompt": "A cinematic scene with clearly directed subject action, camera movement, lighting, pacing, and atmosphere",
  "resolution": "480p"
}

Technical Specs

SpecValue
Model IDalibaba/wan/3.0/image-to-video
Input fieldsprompt (string)<br>image (string)<br>end_image (string)<br>audio (boolean)<br>resolution (string; 480p, 720p, 1080p)<br>duration (integer; 2–30)<br>enable_thinking (boolean)<br>aspect_ratio (string; 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16)<br>seed (integer; 0–2147483647)
Required inputprompt, image
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

Related Models

alibaba/wan/3.0/prime/image-to-videoWan 3.0 Prime by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/3.0/prime/reference-to-videoWan 3.0 Prime by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/3.0/prime/text-to-videoWan 3.0 Prime is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/3.0/reference-to-videoWan 3.0 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/3.0/text-to-videoWan 3.0 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/3.0/videoAlibaba Wan 3.0 Video generates videos up to 30 seconds from text and unified image, video, audio, file, or link media inputs.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.