SandBase is live — $1 in free credits on signupStart free ›

Alibaba modelsvideo generation api

alibaba/wan/2.7/text-to-video

Alibaba Wan 2.7 text-to-video model with cinematic visuals, native audio generation, and configurable duration and resolution.

Input
The text prompt to generate a video from.
The duration of the generated video in seconds. Allowed values: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
The URL of an audio file to guide video generation.
The aspect ratio of the generated video. Allowed values: 16:9, 9:16, 1:1, 4:3, 3:4.
The resolution of the video to generate. Allowed values: 720p, 1080p.
Idle

Example output — click Run to generate your own

API README

Wan 2.7 Text-to-Video

Wan 2.7 Text-to-Video creates a complete short clip from one written scene. The prompt establishes subject, action, camera, light, atmosphere, and style without requiring a starting image. An optional audio file can guide how the visual sequence relates to sound.

The route covers durations from 2 to 15 seconds, five frame shapes, and 720p or 1080p delivery. These choices support quick drafts and higher-resolution final assets through the same text-led workflow. The result is returned as a video URL with file metadata when available.

Highlights

  • Enhanced motion smoothness. Wan 2.7 is documented with improved motion smoothness. This capability applies to action and camera movement generated from text.
  • Superior scene fidelity. The exact text-to-video route highlights stronger fidelity to the described scene. It supports detailed visual direction without an image input.
  • Greater visual coherence. Wan 2.7 is documented to improve visual coherence. The generated scene is intended to remain connected as action develops over time.
  • Natural-language multi-shot generation. The exact route documents multi-shot video controlled through natural language in the prompt. Shot changes are expressed inside the creative brief rather than a separate storyboard field.

Pricing

DurationResolutionPrice
2 seconds720p$0.20
2 seconds1080p$0.30
3 seconds720p$0.30
3 seconds1080p$0.45
4 seconds720p$0.40
4 seconds1080p$0.60
5 seconds720p$0.50
5 seconds1080p$0.75
6 seconds720p$0.60
6 seconds1080p$0.90
7 seconds720p$0.70
7 seconds1080p$1.05
8 seconds720p$0.80
8 seconds1080p$1.20
9 seconds720p$0.90
9 seconds1080p$1.35
10 seconds720p$1.00
10 seconds1080p$1.50
11 seconds720p$1.10
11 seconds1080p$1.65
12 seconds720p$1.20
12 seconds1080p$1.80
13 seconds720p$1.30
13 seconds1080p$1.95
14 seconds720p$1.40
14 seconds1080p$2.10
15 seconds720p$1.50
15 seconds1080p$2.25

When to Use

✅ Good fit❌ Consider alternatives
Produce a food, beauty, or product visual from a written concept.Preserve a specific opening composition; use image-to-video.
Build a short scene whose camera and action are planned entirely in text.Need a single request longer than 15 seconds.
Describe several shots inside one natural-language brief.Need separately structured shot objects in the request.
Time a visual sequence against an existing audio track.Need sound generated without supplying audio.
Prepare landscape, vertical, square, or 4:3 delivery.Need a frame shape outside the documented set.

Prompt Guide

Write the scene in temporal order. Start with setting and subject, then describe action, camera, lighting, atmosphere, and any shot changes; explain how supplied audio should guide timing when used.

Scene: [place, time, atmosphere]
Subject: [appearance and action]
Camera: [framing and movement]
Lighting: [direction, color, progression]
Shots: [natural-language sequence when needed]
Audio relationship: [timing or mood from the supplied track]
{
  "prompt": "A young woman sits beside a rain-streaked café window, then slowly turns toward the camera with a gentle smile as warm interior light reflects in the glass.",
  "duration": 5,
  "resolution": "1080p",
  "aspect_ratio": "16:9"
}

Technical Specs

SpecValue
Required inputprompt
Prompt length3–50,000 characters
Duration2–15 seconds; default 5
Resolution720p default, 1080p
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4
Audio guidanceOptional audio URL
OutputVideo URL with optional file metadata

Related

Related Models