SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/2.7/image-to-video

Alibaba Wan 2.7 image-to-video model transforming still images into video with native audio generation.

Input
The text prompt to generate a video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the starting image for video generation.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the ending image for video generation.
The URL of an audio file to guide video generation.
The duration of the generated video in seconds. Allowed values: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
The aspect ratio of the generated video. Allowed values: 16:9, 9:16, 1:1, 4:3, 3:4.
The resolution of the video to generate. Allowed values: 720p, 1080p.
Idle

Example output — click Run to generate your own

API README

Wan 2.7 Image-to-Video

Wan 2.7 Image-to-Video turns an existing visual into a configurable short clip. A starting image and written direction form the core request, with optional ending-frame and audio guidance available when the transition needs additional anchors. The workflow returns a video URL ready for downstream playback or delivery.

Teams can choose between 720p and 1080p, five frame shapes, and durations from 2 to 15 seconds. These choices let one route cover quick drafts as well as higher-resolution deliverables. The model is positioned for image-led work where the original scene must remain the basis of the clip.

Highlights

  • Enhanced motion smoothness. Wan 2.7 is documented with improved motion smoothness. This capability applies to motion generated from the supplied image.
  • Superior scene fidelity. The exact image-to-video route highlights stronger fidelity to the generated scene. This is a model capability rather than a resolution setting.
  • Greater visual coherence. Wan 2.7 is documented to improve visual coherence. The result is intended to remain visually connected as the scene changes over time.
  • Audio-driven animation. Wan 2.7 supports audio-driven image-to-video generation. Supplied sound can guide how movement unfolds from the opening frame.

Pricing

DurationResolutionPrice
2 seconds720p$0.20
2 seconds1080p$0.30
3 seconds720p$0.30
3 seconds1080p$0.45
4 seconds720p$0.40
4 seconds1080p$0.60
5 seconds720p$0.50
5 seconds1080p$0.75
6 seconds720p$0.60
6 seconds1080p$0.90
7 seconds720p$0.70
7 seconds1080p$1.05
8 seconds720p$0.80
8 seconds1080p$1.20
9 seconds720p$0.90
9 seconds1080p$1.35
10 seconds720p$1.00
10 seconds1080p$1.50
11 seconds720p$1.10
11 seconds1080p$1.65
12 seconds720p$1.20
12 seconds1080p$1.80
13 seconds720p$1.30
13 seconds1080p$1.95
14 seconds720p$1.40
14 seconds1080p$2.10
15 seconds720p$1.50
15 seconds1080p$2.25

When to Use

✅ Good fit❌ Consider alternatives
Turn a product photograph into a short promotional clip.Create the whole scene without an existing image; use text-to-video.
Add life to an illustration for a social post or display.Require a single request longer than 15 seconds.
Build a before-to-after transition from two prepared frames.Need output above the documented 1080p tier.
Time visual changes against a supplied audio track.Need sound generated without an audio input.
Prepare one asset in landscape, portrait, square, or 4:3 delivery.Need a frame shape outside the five documented choices.

Prompt Guide

Describe movement relative to the supplied image. State the subject action, camera path, environmental motion, and final state; add an audio URL only when sound guidance is available.

Subject action: [movement over time]
Camera: [pan, follow, orbit, push, or static]
Environment: [weather, light, background motion]
Timing: [pace and progression]
Final state: [desired arrival at the optional end image]
Audio relationship: [how supplied audio should guide the clip]
{
  "image": "<start-image-url>",
  "prompt": "The bus starts moving slowly down the street. Rain starts, pedestrians open umbrellas, and the camera follows the bus.",
  "duration": 5,
  "resolution": "1080p",
  "aspect_ratio": "16:9"
}

Technical Specs

SpecValue
Required inputsimage, prompt
Prompt length3–50,000 characters
Duration2–15 seconds; default 5
Resolution720p default, 1080p
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4
Frame guidanceRequired image; optional end_image
Audio guidanceOptional audio URL
OutputVideo URL with optional file metadata

Related

Related Models