meituan/longcat-multi-avatar/image-audio-to-video
Longcat Multi Avatar Image Audio To Video by meituan - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
PNG, JPEG, WebP, or GIF · 20 MiB maximum
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runmeituan/longcat-multi-avatar/image-audio-to-videoInput Schema
13 parameters · 2 required · 11 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
image | string | Required | The URL of the image containing two speakers. |
prompt | string | Required | The prompt to guide the video generation. · Default: "Two people are having a conversation with natural expressions and movements." |
seed | integer | Optional | The seed for the random number generator. |
audio_type | string | Optional | How to combine the two audio tracks. 'para' (parallel) plays both simultaneously, 'add' (sequential) plays person 1 first then person 2. · Options: para, add · Default: "para" paraadd |
resolution | string | Optional | Resolution of the generated video (480p or 720p). Billing is per video-second (16 frames): 480p is 1 unit per second and 720p is 4 units per second. · Options: 480p, 720p · Default: "480p" 480p720p |
bbox_person1 | string | Optional | Bounding box for person 1. If not provided, defaults to left half of image. |
bbox_person2 | string | Optional | Bounding box for person 2. If not provided, defaults to right half of image. |
num_segments | integer | Optional | Number of video segments to generate. Each segment adds ~5 seconds of video. First segment is ~5.8s, additional segments are 5s each. · Min: 1 · Max: 10 · Default: 1 |
audio_url_person1 | string | Optional | The URL of the audio file for person 1 (left side). · Default: "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_man.WAV" |
audio_url_person2 | string | Optional | The URL of the audio file for person 2 (right side). · Default: "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_woman.WAV" |
num_inference_steps | integer | Optional | The number of inference steps to use. · Min: 10 · Max: 100 · Default: 30 |
text_guidance_scale | number | Optional | The text guidance scale for classifier-free guidance. · Min: 1 · Max: 10 · Default: 4 |
audio_guidance_scale | number | Optional | The audio guidance scale. Higher values may lead to exaggerated mouth movements. · Min: 1 · Max: 10 · Default: 4 |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "meituan/longcat-multi-avatar/image-audio-to-video",
"image": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing.png",
"prompt": "Static camera, In a professional recording studio, two people stand facing each other, both wearing large headphones. They are speaking clearly into a large condenser microphone suspended between them. They looked at each other affectionately and occasionally shook their heads according to the rhythm. The soundproofed walls and visible recording equipment create an atmosphere focused on capturing high-quality audio as they interact and communicate.",
"audio_type": "para",
"resolution": "480p",
"num_segments": 1,
"audio_url_person1": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_man.WAV",
"audio_url_person2": "https://raw.githubusercontent.com/meituan-longcat/LongCat-Video/refs/heads/main/assets/avatar/multi/sing_woman.WAV",
"num_inference_steps": 30,
"text_guidance_scale": 4,
"audio_guidance_scale": 4
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Longcat Multi Avatar
Longcat Multi Avatar is a LongCat route built for image-and-audio-driven avatar animation. It turns a portrait reference and speech audio into a synchronized talking-character video, giving creators a focused endpoint instead of forcing one generic workflow across materially different production tasks. The route belongs to a family known for efficient generation, stable visual continuity, and practical media workflows, so it is best evaluated as a creative system for intentional shots and assets rather than as a one-click novelty generator.
Use this endpoint when the input contract and deliverable match that job exactly. Its documented controls include seed (The seed for the random number generator); audio_type choices (para, add); resolution choices (480p, 720p); bbox_person1 (Bounding box for person 1. If not provided, defaults to left half of image). Together, these controls help teams plan predictable iterations, compare outputs under stable settings, and connect generation to high-volume creative iteration and avatar or video production without hiding the operational choices that shape the result.
Highlights
- Purpose-built Image-and-audio-driven avatar animation. The route accepts a portrait reference and speech audio and produces a synchronized talking-character video; its interface is scoped to that transformation, keeping source assets and creative intent explicit.
- Creative direction. Prompts can describe subject behavior, composition, camera intent, lighting, material, atmosphere, and temporal progression so the result is driven by a shot plan rather than isolated keywords.
- Route-specific control. The request exposes seed (The seed for the random number generator); audio_type choices (para, add); resolution choices (480p, 720p); bbox_person1 (Bounding box for person 1. If not provided, defaults to left half of image), allowing the same concept to be tested systematically while preserving a repeatable production setup.
- Pipeline-ready output. The generated media asset is returned through the documented asynchronous output contract, which suits review queues, batch iteration, and downstream automation. Editors can review pacing, continuity, lens language, choreography, transitions, temporal artifacts, soundtrack alignment, color response, delivery framing, and cut compatibility before approval.
Pricing
| Configuration | Price |
|---|---|
| Billing rule | params.resolution == "480p" ? params.duration * 0.15 : params.duration * 0.30 |
| resolution=480p | Calculated by billing rule |
| resolution=720p | Calculated by billing rule |
When to Use
| Scenario | Why this model fits |
|---|---|
| Create the exact route output | Choose it when you need image-and-audio-driven avatar animation and already have a portrait reference and speech audio. |
| Develop controlled variations | Keep the main brief fixed while changing one documented setting at a time to compare motion, framing, quality, or asset behavior. |
| Build repeatable batches | Use a consistent request shape for catalog, campaign, storyboard, game-asset, or social-content production. |
| Preserve source intent | Prefer this route when the supplied reference material must remain the foundation of a synchronized talking-character video. |
| Connect a media pipeline | Use asynchronous results in an automated review, approval, post-production, or asset-management workflow. |
Prompt Guide
Start with the desired result, then describe the source relationship, subject action, composition or camera behavior, lighting, style, and timing. For image-and-audio-driven avatar animation, state what must remain stable as clearly as what should change. Use only fields exposed by the schema; the example below is structurally valid for this route.
{
"prompt": "Static camera, In a professional recording studio, two people stand facing each other, both wearing large headphones. They are speaking clearly into a large condenser microphone suspended between them. They looked at each other affectionately and occasionally shook their heads according to the rhythm. The soundproofed walls and visible recording equipment create an atmosphere focused on capturing high-quality audio as they interact and communicate.",
"image": "https://example.com/image.jpg",
"seed": 1,
"audio_type": "para",
"resolution": "480p"
}
Technical Specs
| Specification | Value |
|---|---|
| Model ID | meituan/longcat-multi-avatar/image-audio-to-video |
| Workflow | Image-and-audio-driven avatar animation |
| Required inputs | prompt, image |
seed | integer |
image | string |
prompt | string |
audio_type | string; options: para, add |
resolution | string; options: 480p, 720p |
bbox_person1 | string |
bbox_person2 | string |
num_segments | integer; minimum: 1; maximum: 10 |
audio_url_person1 | string |
audio_url_person2 | string |

