Skip to content

Wan 2.2 Speech to Video

POST/v1/run

Wan 2.2 Speech To Video by Alibaba - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Request body

Submit an async generation request. The model field selects the model; other fields are model-specific input parameters.

stringmodelrequired

Model identifier. Set to alibaba/wan/2.2/speech-to-video.

Default: alibaba/wan/2.2/speech-to-video

stringpromptrequired

The text prompt used for video generation.

stringimagerequired

URL of the input image. If the input image does not match the chosen aspect ratio, it is resized and center cropped.

Optional<string>audio

The URL of the audio file.

Optional<string>resolution

Resolution of the generated video (480p, 580p, or 720p).

Allowed values: 480p, 720p

Default: 480p

Optional<integer>seed

Random seed for reproducibility. If None, a random seed is chosen.

Response Schema

The submit endpoint returns an accepted generation task. Poll the result endpoint with the returned id for terminal outputs or errors.

Optional<string>error

Error message if the task failed. Empty on success.

stringidrequired

Unique identifier for the generation task.

Optional<string>model

Model ID used for the prediction.

Optional<array>outputs

Array of generated content. Empty when status is not completed.

stringstatusrequired

Status of the task: pending, running, completed, failed, or timeout.

Allowed values: pending, running, completed, failed, timeout

Model capabilities

array<string>capability_tagsrequired

Capabilities declared by the model registry.

Default: audio-to-video

stringexecution_moderequired

Execution mode declared by the model registry.

Default: async