SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/3.0/image-to-video

Wan 3.0 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.

Input
Text prompt describing the motion to generate.
Include generated audio.
Output video resolution tier. Allowed values: 480p, 720p, 1080p.
230
Output duration in seconds. Set to null for smart duration, which lets the model pick a length from the prompt and reference media. Range: 2 to 30.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
02147483647
Range: 0 to 2147483647.
Enter a JSON array.
Idle

Example output — click Run to generate your own

API README

Alibaba Wan 3.0 Image-to-Video

Animate a supplied still image into a video. Use the image as the visual starting point, then describe motion, camera behavior, atmosphere, and optional sound in the prompt.

Good fit

  • Bringing product photos, illustrations, or portraits to life
  • Controlled first-frame or last-frame transitions
  • Short cinematic loops and motion tests

Input guide

prompt and at least one media item are required. Each media entry is a typed object {type, url}: use first_frame or last_frame for boundary images, and reference_image for appearance guidance. Use an accessible media URL in url. media is a required array with at least 1 item. Each item is an object {type, url}; type is one of first_frame, last_frame, reference_image, reference_video, reference_audio, file, or link. Set duration from 2–30 seconds (default 5), choose 480p/720p/1080p (default 1080p), and use aspect_ratio, seed, or audio as needed. Write the desired movement rather than restating every pixel of the source image.

Example prompt:

The subject slowly turns toward the camera as a light breeze moves the fabric; subtle handheld push-in, natural daylight, no dialogue. prompt is a string and required, with a maximum of 5,000 characters. resolution is an enum: 480p, 720p, or 1080p (default 1080p). duration is an integer from 2 to 30 (default 5). aspect_ratio is an enum: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, or 9:16. audio is a boolean (default true). seed is optional and must be an integer from 0 to 2,147,483,647.

Output and pricing

The asynchronous response returns a downloadable video URL and may include its content type. Standard Wan 3.0 is billed per output second: $0.05 at 480p, $0.10 at 720p, and $0.20 at 1080p. A default 5-second 1080p render costs $1.00. Audio and aspect ratio do not add a separate charge.

Related Models

alibaba/wan/3.0/prime/videoWan 3.0 Prime by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion and camera movement.alibaba/wan/3.0/videoAlibaba Wan 3.0 Video generates videos up to 30 seconds from text and unified image, video, audio, file, or link media inputs.alibaba/wan/2.1/image-to-video/loraWan 2.1 Lora is Alibaba's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.alibaba/wan/2.1/text-to-videoWan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/2.1/text-to-video/loraWan 2.1 Lora by Alibaba - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.alibaba/wan/2.1/vaceWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.