SandBase is live — $1 in free credits on signupStart free ›

meituan modelsvideo generation api

meituan/longcat-single-avatar/image-audio-to-video

Longcat Single Avatar Image Audio To Video by sandbase-ai - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Input
The prompt to guide the video generation.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the image to animate.
The URL of the audio file to drive the avatar.
Resolution of the generated video (480p or 720p). Billing is per video-second (16 frames): 480p is 1 unit per second and 720p is 4 units per second. Allowed values: 480p, 720p.
110
The audio guidance scale. Higher values may lead to exaggerated mouth movements. Range: 1 to 10.
110
The text guidance scale for classifier-free guidance. Range: 1 to 10.
110
Number of video segments to generate. Each segment adds ~5 seconds of video. First segment is ~5.8s, additional segments are 5s each. Range: 1 to 10.
The seed for the random number generator.
10100
Range: 10 to 100.
Idle

Example output — click Run to generate your own

API README

LongCat Single Avatar

LongCat Single Avatar is a LongCat route built for image-and-audio-driven avatar animation. It turns a portrait reference and speech audio into a synchronized talking-character video, giving creators a focused endpoint instead of forcing one generic workflow across materially different production tasks. The route belongs to a family known for efficient generation, stable visual continuity, and practical media workflows, so it is best evaluated as a creative system for intentional shots and assets rather than as a one-click novelty generator.

Use this endpoint when the input contract and deliverable match that job exactly. Its documented controls include seed (The seed for the random number generator); resolution choices (480p, 720p); num_segments (Number of video segments to generate. Each segment adds ~5 seconds of video. First segment is ~5.8s, additional segments are 5s each); text_guidance_scale (The text guidance scale for classifier-free guidance). Together, these controls help teams plan predictable iterations, compare outputs under stable settings, and connect generation to high-volume creative iteration and avatar or video production without hiding the operational choices that shape the result.

Highlights

  • Purpose-built Image-and-audio-driven avatar animation. The route accepts a portrait reference and speech audio and produces a synchronized talking-character video; its interface is scoped to that transformation, keeping source assets and creative intent explicit.
  • Creative direction. Prompts can describe subject behavior, composition, camera intent, lighting, material, atmosphere, and temporal progression so the result is driven by a shot plan rather than isolated keywords.
  • Route-specific control. The request exposes seed (The seed for the random number generator); resolution choices (480p, 720p); num_segments (Number of video segments to generate. Each segment adds ~5 seconds of video. First segment is ~5.8s, additional segments are 5s each); text_guidance_scale (The text guidance scale for classifier-free guidance), allowing the same concept to be tested systematically while preserving a repeatable production setup.
  • Pipeline-ready output. The generated media asset is returned through the documented asynchronous output contract, which suits review queues, batch iteration, and downstream automation. Editors can review pacing, continuity, lens language, choreography, transitions, temporal artifacts, soundtrack alignment, color response, delivery framing, and cut compatibility before approval.

Pricing

ConfigurationPrice
Billing ruleparams.resolution == "480p" ? params.duration * 0.15 : params.duration * 0.30
resolution=480pCalculated by billing rule
resolution=720pCalculated by billing rule

When to Use

ScenarioWhy this model fits
Create the exact route outputChoose it when you need image-and-audio-driven avatar animation and already have a portrait reference and speech audio.
Develop controlled variationsKeep the main brief fixed while changing one documented setting at a time to compare motion, framing, quality, or asset behavior.
Build repeatable batchesUse a consistent request shape for catalog, campaign, storyboard, game-asset, or social-content production.
Preserve source intentPrefer this route when the supplied reference material must remain the foundation of a synchronized talking-character video.
Connect a media pipelineUse asynchronous results in an automated review, approval, post-production, or asset-management workflow.

Prompt Guide

Start with the desired result, then describe the source relationship, subject action, composition or camera behavior, lighting, style, and timing. For image-and-audio-driven avatar animation, state what must remain stable as clearly as what should change. Use only fields exposed by the schema; the example below is structurally valid for this route.

{
  "prompt": "A western man stands on stage under dramatic lighting, holding a microphone close to their mouth. Wearing a vibrant red jacket with gold embroidery, the singer is speaking while smoke swirls around them, creating a dynamic and atmospheric scene.",
  "image": "https://example.com/image.jpg",
  "seed": 1,
  "audio": "https://example.com/audio.mp3",
  "resolution": "480p"
}

Technical Specs

SpecificationValue
Model IDmeituan/longcat-single-avatar/image-audio-to-video
WorkflowImage-and-audio-driven avatar animation
Required inputsprompt, image
seedinteger
audiostring
imagestring
promptstring
resolutionstring; options: 480p, 720p
num_segmentsinteger; minimum: 1; maximum: 10
num_inference_stepsinteger; minimum: 10; maximum: 100
text_guidance_scalenumber; minimum: 1; maximum: 10
audio_guidance_scalenumber; minimum: 1; maximum: 10

Related Models

Related Models

meituan/longcat-video/distilled/image-to-videoLongcat Video Distilled is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/distilled/image-to-video-480pLongcat Video Distilled Image To Video 480p is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/distilled/text-to-video-720pLongcat Video Distilled Text To Video 720p by sandbase-ai - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.meituan/longcat-video/distilled/text-to-video/480pLongcat Video Distilled 480p by meituan - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.meituan/longcat-video/image-to-videoLongcat Video is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/image-to-video-480pLongcat Video Image To Video 480p is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/text-to-video/480pLongcat Video 480p by meituan - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.meituan/longcat-video/text-to-video/720pLongcat Video 720p by meituan - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.