Alibaba modelsvideo generation api

alibaba/wan/3.0/prime/reference-to-video

Wan 3.0 Prime by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.

Input
Text prompt directing how the reference media is used. Reference media can be addressed positionally, e.g. 'the subject in Image 1 walks past Video 1'.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
Include generated audio.
Output video resolution tier. Allowed values: 480p, 720p, 1080p.
230
Output duration in seconds. Set to null for smart duration, which lets the model pick a length from the prompt and reference media. Range: 2 to 30.
02147483647
Range: 0 to 2147483647.
Up to 10 reference image URLs.
Up to 5 reference audio URLs totaling at most 15 seconds.
Up to 5 reference video URLs totaling at most 15 seconds. Each clip must be at least 16 fps.
Document URL to base the video on. Requires enable_thinking=true.
Public webpage URL to base the video on. Requires enable_thinking=true. Only pages that do not require login can be read.
Enable enhanced reasoning before generation.
Idle

Example output — click Run to generate your own

API README

Wan 3.0 Prime

Wan 3.0 Prime uses one or more subject references to produce a new sequence with the latest 3.0 family’s emphasis on coherent motion, scene detail, and cinematic presentation. Its creative focus is carrying recognizable subjects into new actions and viewpoints, with the Prime configuration favoring greater visual refinement and production polish.

Plan the clip around a clear temporal arc: define the initial state, meaningful action, camera behavior, environmental response, and intended ending. Identify each referenced subject and explain its role in the scene. Reserve Prime for material where fine detail and finishing quality justify the higher rate.

Highlights

  • 3.0 subject-reference grounding. Carries referenced characters or objects into a newly staged sequence.
  • Improved motion coherence. Maintains more stable subject action and environmental dynamics across the generated clip.
  • Prime visual refinement. Prioritizes fine textures, lighting, composition, and cleaner production detail.
  • Viewpoint-flexible identity. Keeps recognizable traits as referenced subjects change pose, action, and camera angle.

Pricing

ConfigurationBilling unitPrice
Generated duration480p, per second$0.075
Generated duration720p, per second$0.150
Generated duration1080p, per second$0.300

When to Use

✅ Good fit❌ Consider alternatives
The project needs this exact reference-to-video workflowThe intended task belongs to a different media route
Available source media matches every required fieldRequired assets or usage rights are unavailable
The brief can define subject, action, camera, and styleOutput must be deterministic at frame or pixel level
Supported duration, resolution, and framing fit deliveryFinal placement requires unsupported specifications
An asynchronous generation job fits productionA live or frame-synchronous response is mandatory

Prompt Guide

Describe the result as a shot or design brief: identify subjects and references, state the action or transformation, specify environment and composition, then add camera behavior, lighting, pacing, style, sound, and preservation constraints where relevant. Use exact reference identifiers exposed by the local schema.

{
  "aspect_ratio": "21:9",
  "audio": true,
  "prompt": "A cinematic scene with clearly directed subject action, camera movement, lighting, pacing, and atmosphere",
  "resolution": "480p"
}

Technical Specs

SpecValue
Model IDalibaba/wan/3.0/prime/reference-to-video
Input fieldsprompt (string)<br>aspect_ratio (string; 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16)<br>audio (boolean)<br>resolution (string; 480p, 720p, 1080p)<br>duration (integer; 2–30)<br>seed (integer; 0–2147483647)<br>reference_image_urls (array)<br>reference_audio_urls (array)<br>reference_video_urls (array)<br>file_url (string)<br>web_url (string)<br>enable_thinking (boolean)
Required inputprompt
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

Related Models

alibaba/wan/3.0/prime/image-to-videoWan 3.0 Prime by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/3.0/prime/text-to-videoWan 3.0 Prime is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/3.0/image-to-videoWan 3.0 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/3.0/reference-to-videoWan 3.0 by Alibaba - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.alibaba/wan/3.0/text-to-videoWan 3.0 is Alibaba's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.alibaba/wan/3.0/videoAlibaba Wan 3.0 Video generates videos up to 30 seconds from text and unified image, video, audio, file, or link media inputs.alibaba/wan/2.1/vace/depthWan 2.1 Vace by Alibaba - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.alibaba/wan/2.1/vace/inpaintingWan 2.1 Vace is Alibaba's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.