SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/2.5/image-to-video

Wan 2.5 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.

Input
The text prompt describing the desired video motion. Max 800 characters.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the image to use as the first frame. Must be publicly accessible or base64 data URI.
URL of the audio to use as the background music. Must be publicly accessible. Limit handling: If the audio duration exceeds the duration value (5 or 10 seconds), the audio is truncated to the first 5 or 10 seconds, and the rest is discarded. If the audio is shorter than the video, the remaining part of the video will be silent. For example, if the audio is 3 seconds long and the video duration is 5 seconds, the first 3 seconds of the output video will have sound, and the last 2 seconds will be silent. - Format: WAV, MP3. - Duration: 3 to 30 s. - File size: Up to 15 MB.
Video resolution. Valid values: 480p, 720p, 1080p Allowed values: 480p, 720p, 1080p.
Duration of the generated video in seconds. Choose between 5 or 10 seconds. Allowed values: 5, 10.
Random seed for reproducibility. If None, a random seed is chosen.
Idle

Example output — click Run to generate your own

API README

Wan 2.5 Image to Video

Wan 2.5 Image to Video animates one still image into a short generated sequence. The image establishes the opening visual state, while the prompt describes how the subject, camera, and environment should move. It is a focused path from a selected frame to a motion concept.

The route produces five- or ten-second clips at three delivery resolutions. Optional background audio can accompany the generated video when a soundtrack is already available. The result is a new video rather than an edit of an existing clip.

Highlights

  • First-frame image animation. The exact workflow uses the supplied image as the first frame. This gives the generated motion a concrete visual starting point.
  • Text-directed motion generation. A prompt describes how the still scene should change over time. Subject action and camera behavior can be expressed together.
  • Strong prompt adherence. Wan 2.5 is documented with substantially improved prompt adherence over Wan 2.2. It follows detailed creative instructions spanning action, camera movement, and stylistic direction with greater accuracy.
  • Improved visual quality. Wan 2.5 advances the image quality of the 2.5 image-to-video model, giving generated frames a more refined visual finish than the preceding generation.

Pricing

DurationResolutionPrice
5 seconds480p$0.25
5 seconds720p$0.50
5 seconds1080p$0.75
10 seconds480p$0.50
10 seconds720p$1.00
10 seconds1080p$1.50

When to Use

✅ Good fit❌ Consider alternatives
Animating a selected hero frame into a short revealGenerating a scene with no opening image
Adding camera movement to a still conceptEditing an existing video timeline
Creating five- or ten-second social clipsRequiring another duration
Delivering the idea at one of three resolution tiersRequiring 4K output
Pairing the result with an available music bedGenerating lip-synchronized dialogue

Prompt Guide

Describe motion relative to the opening image. Separate subject action, camera movement, environmental movement, pacing, and final state.

Subject action: [movement]
Camera: [framing and movement]
Environment: [background motion]
Pacing: [tempo]
Final state: [shot conclusion]
{
  "image": "https://example.com/opening-frame.jpg",
  "prompt": "The cyclist accelerates along the ridge as the camera tracks from the side, ending on a wide valley view.",
  "duration": 5,
  "resolution": "1080p",
  "seed": 42
}

Technical Specs

SpecValue
Model IDalibaba/wan/2.5/image-to-video
Required inputsimage and prompt
Duration5 or 10 seconds; default 5
Resolution480p, 720p, or 1080p; default 1080p
Optional audioPublic WAV or MP3 URL, 3–30 seconds, up to 15 MB
RepeatabilityOptional integer seed
OutputVideo URL and content type
ExecutionAsynchronous job

Related

Related Models

alibaba/wan/2.5/text-to-videoWan 2.5 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/2.1/image-to-videoWan 2.1 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/2.1/image-to-video/loraWan 2.1 Lora is Alibaba's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.alibaba/wan/2.1/text-to-videoWan 2.1 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/2.1/text-to-video/loraWan 2.1 Lora by Alibaba - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.alibaba/wan/2.1/vaceWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.