API reference · meituan

meituan/longcat-multi-avatar/image-audio-to-video

Integrate this model through SandBase's unified API, with production-ready schemas and examples.

IMAGEAsyncOpen model
Production endpoint

Send your first request

OpenAI-compatible endpoint with unified authentication and usage tracking.

POSThttps://api.sandbase.ai/v1/run
Model IDmeituan/longcat-multi-avatar/image-audio-to-video
01

Input Schema

13 parameters · 2 required · 11 optional

ParameterTypeRequiredDescription
imagestringRequiredThe URL of the image containing two speakers.
promptstringRequiredThe prompt to guide the video generation. · Default: "Two people are having a conversation with natural expressions and movements."
seedintegerOptionalThe seed for the random number generator.
audio_typestringOptionalHow to combine the two audio tracks. 'para' (parallel) plays both simultaneously, 'add' (sequential) plays person 1 first then person 2. · Options: para, add · Default: "para"
paraadd
resolutionstringOptionalResolution of the generated video (480p or 720p). Billing is per video-second (16 frames): 480p is 1 unit per second and 720p is 4 units per second. · Options: 480p, 720p · Default: "480p"
480p720p
bbox_person1stringOptionalBounding box for person 1. If not provided, defaults to left half of image.
bbox_person2stringOptionalBounding box for person 2. If not provided, defaults to right half of image.
num_segmentsintegerOptionalNumber of video segments to generate. Each segment adds ~5 seconds of video. First segment is ~5.8s, additional segments are 5s each. · Min: 1 · Max: 10 · Default: 1
audio_url_person1stringOptionalThe URL of the audio file for person 1 (left side). · Default: "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_man.WAV"
audio_url_person2stringOptionalThe URL of the audio file for person 2 (right side). · Default: "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_woman.WAV"
num_inference_stepsintegerOptionalThe number of inference steps to use. · Min: 10 · Max: 100 · Default: 30
text_guidance_scalenumberOptionalThe text guidance scale for classifier-free guidance. · Min: 1 · Max: 10 · Default: 4
audio_guidance_scalenumberOptionalThe audio guidance scale. Higher values may lead to exaggerated mouth movements. · Min: 1 · Max: 10 · Default: 4
02

Output Schema

FieldTypeDescription
idstringUnique identifier for the generation task
statusstringTask status: pending, running, completed, failed, timeout
modelstringModel used for the generation
outputsarrayArray of output items
outputs[].urlstringURL of the generated artifact
outputs[].content_typestringMIME type (e.g. image/png, video/mp4)
errorobject | nullError details if failed, null on success
error.typestringMachine-readable error type code
error.messagestringHuman-readable error description

Async Workflow

This model uses asynchronous execution. Submit a request and poll for the result.

  1. Submit — POST to /v1/run, receive an id
  2. Poll — GET /v1/run/{id} until status is completed, failed, or timeout
  3. Retrieve — Read outputs from the completed response
03

Code Examples

Ready-to-run snippets

# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
  "model": "meituan/longcat-multi-avatar/image-audio-to-video",
  "image": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing.png",
  "prompt": "Static camera, In a professional recording studio, two people stand facing each other, both wearing large headphones. They are speaking clearly into a large condenser microphone suspended between them. They looked at each other affectionately and occasionally shook their heads according to the rhythm. The soundproofed walls and visible recording equipment create an atmosphere focused on capturing high-quality audio as they interact and communicate.",
  "audio_type": "para",
  "resolution": "480p",
  "num_segments": 1,
  "audio_url_person1": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_man.WAV",
  "audio_url_person2": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_woman.WAV",
  "num_inference_steps": 30,
  "text_guidance_scale": 4,
  "audio_guidance_scale": 4
}'

# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
  -H "Authorization: Bearer YOUR_API_KEY"