SandBase is live — $1 in free credits on signupStart free ›

meituan modelsvideo generation api

meituan/longcat-video/distilled/text-to-video-720p

Longcat Video Distilled Text To Video 720p by sandbase-ai - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.

Input
The prompt to guide the video generation.
The aspect ratio of the generated image. Allowed values: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16.
The seed for the random number generator.
216
The number of inference steps to use for refinement. Range: 2 to 16.
216
Range: 2 to 16.
Idle

Example output — click Run to generate your own

API README

LongCat Video Distilled

LongCat Video Distilled Text to Video 720p converts a written scene directly into an HD-ready draft without a reference image. Its distilled path balances economical iteration with a 720p canvas that is large enough for composition, motion, and edit-timing review. It suits storyboard sequences, pitch-film drafts, and campaign exploration when the creative direction is still evolving but reviewers need more visual information than a proxy-resolution render provides.

The 720p route adds refinement-step control alongside fps, frame count, aspect ratio, guidance, quality, write mode, and output-format selection. That extra refinement stage distinguishes it from the leaner 480p distilled option and supports closer examination of promising shots. Write the prompt as a sequence of visible beats, reserve fine texture for later passes, and compare refinement settings while leaving camera and subject direction unchanged.

Highlights

  • Purpose-built Text-to-video generation. The route accepts a shot description with subject, action, camera, and atmosphere and produces a coherent video interpretation of the written direction; its interface is scoped to that transformation, keeping source assets and creative intent explicit.
  • Creative direction. Prompts can describe subject behavior, composition, camera intent, lighting, material, atmosphere, and temporal progression so the result is driven by a shot plan rather than isolated keywords.
  • Route-specific control. The request exposes seed (The seed for the random number generator); aspect_ratio choices (21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16); num_refine_inference_steps (The number of inference steps to use for refinement), allowing the same concept to be tested systematically while preserving a repeatable production setup.
  • Pipeline-ready output. The generated media asset is returned through the documented asynchronous output contract, which suits review queues, batch iteration, and downstream automation. Editors can review pacing, continuity, lens language, choreography, transitions, temporal artifacts, soundtrack alignment, color response, delivery framing, and cut compatibility before approval.

Pricing

ConfigurationPrice
Billing ruleparams.duration * 0.01
Formula-priced requestStarts from $0.010000; final charge follows the billing rule above

When to Use

ScenarioWhy this model fits
Create the exact route outputChoose it when you need text-to-video generation and already have a shot description with subject, action, camera, and atmosphere.
Develop controlled variationsKeep the main brief fixed while changing one documented setting at a time to compare motion, framing, quality, or asset behavior.
Build repeatable batchesUse a consistent request shape for catalog, campaign, storyboard, game-asset, or social-content production.
Preserve source intentPrefer this route when the supplied reference material must remain the foundation of a coherent video interpretation of the written direction.
Connect a media pipelineUse asynchronous results in an automated review, approval, post-production, or asset-management workflow.

Prompt Guide

Start with the desired result, then describe the source relationship, subject action, composition or camera behavior, lighting, style, and timing. For text-to-video generation, state what must remain stable as clearly as what should change. Use only fields exposed by the schema; the example below is structurally valid for this route.

{
  "prompt": "realistic filming style, a person wearing a dark helmet, a deep-colored jacket, blue jeans, and bright yellow shoes rides a skateboard along a winding mountain road. The skateboarder starts in a standing position, then gradually lowers into a crouch, extending one hand to touch the road surface while maintaining a low center of gravity to navigate a sharp curve. After completing the turn, the skateboarder rises back to a standing position and continues gliding forward. The background features lush green hills flanking both sides of the road, with distant snow-capped mountain peaks rising against a clear, bright blue sky. The camera follows closely from behind, smoothly tracking the skateboarder’s movements and capturing the dynamic scenery along the route. The scene is shot in natural daylight, highlighting the vivid outdoor environment and the skateboarder’s fluid actions.",
  "seed": 1,
  "aspect_ratio": "21:9",
  "num_inference_steps": 12,
  "num_refine_inference_steps": 12
}

Technical Specs

SpecificationValue
Model IDmeituan/longcat-video/distilled/text-to-video-720p
WorkflowText-to-video generation
Required inputsprompt
seedinteger
promptstring
aspect_ratiostring; options: 21:9, 16:9, 3:2, 4:3, 5:4, 1:1, 4:5, 3:4, 2:3, 9:16
num_inference_stepsinteger; minimum: 2; maximum: 16
num_refine_inference_stepsinteger; minimum: 2; maximum: 16

Related Models

Related Models

meituan/longcat-video/distilled/image-to-videoLongcat Video Distilled is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/distilled/image-to-video-480pLongcat Video Distilled Image To Video 480p is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/distilled/text-to-video/480pLongcat Video Distilled 480p by meituan - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.meituan/longcat-video/image-to-videoLongcat Video is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/image-to-video-480pLongcat Video Image To Video 480p is meituan's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.meituan/longcat-video/text-to-video/480pLongcat Video 480p by meituan - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.meituan/longcat-video/text-to-video/720pLongcat Video 720p by meituan - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.meituan/longcat-single-avatar/image-audio-to-videoLongcat Single Avatar Image Audio To Video by sandbase-ai - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.