API reference · meituan
meituan/longcat-multi-avatar/image-audio-to-video
Integrate this model through SandBase's unified API, with production-ready schemas and examples.
Production endpoint
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
POST
https://api.sandbase.ai/v1/runModel ID
meituan/longcat-multi-avatar/image-audio-to-video01
Input Schema
13 parameters · 2 required · 11 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
image | string | Required | The URL of the image containing two speakers. |
prompt | string | Required | The prompt to guide the video generation. · Default: "Two people are having a conversation with natural expressions and movements." |
seed | integer | Optional | The seed for the random number generator. |
audio_type | string | Optional | How to combine the two audio tracks. 'para' (parallel) plays both simultaneously, 'add' (sequential) plays person 1 first then person 2. · Options: para, add · Default: "para" paraadd |
resolution | string | Optional | Resolution of the generated video (480p or 720p). Billing is per video-second (16 frames): 480p is 1 unit per second and 720p is 4 units per second. · Options: 480p, 720p · Default: "480p" 480p720p |
bbox_person1 | string | Optional | Bounding box for person 1. If not provided, defaults to left half of image. |
bbox_person2 | string | Optional | Bounding box for person 2. If not provided, defaults to right half of image. |
num_segments | integer | Optional | Number of video segments to generate. Each segment adds ~5 seconds of video. First segment is ~5.8s, additional segments are 5s each. · Min: 1 · Max: 10 · Default: 1 |
audio_url_person1 | string | Optional | The URL of the audio file for person 1 (left side). · Default: "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_man.WAV" |
audio_url_person2 | string | Optional | The URL of the audio file for person 2 (right side). · Default: "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_woman.WAV" |
num_inference_steps | integer | Optional | The number of inference steps to use. · Min: 10 · Max: 100 · Default: 30 |
text_guidance_scale | number | Optional | The text guidance scale for classifier-free guidance. · Min: 1 · Max: 10 · Default: 4 |
audio_guidance_scale | number | Optional | The audio guidance scale. Higher values may lead to exaggerated mouth movements. · Min: 1 · Max: 10 · Default: 4 |
02
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
03
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "meituan/longcat-multi-avatar/image-audio-to-video",
"image": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing.png",
"prompt": "Static camera, In a professional recording studio, two people stand facing each other, both wearing large headphones. They are speaking clearly into a large condenser microphone suspended between them. They looked at each other affectionately and occasionally shook their heads according to the rhythm. The soundproofed walls and visible recording equipment create an atmosphere focused on capturing high-quality audio as they interact and communicate.",
"audio_type": "para",
"resolution": "480p",
"num_segments": 1,
"audio_url_person1": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_man.WAV",
"audio_url_person2": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_woman.WAV",
"num_inference_steps": 30,
"text_guidance_scale": 4,
"audio_guidance_scale": 4
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"
