SandBase is live — $1 in free credits on signupStart free ›

Lightricks modelsvideo generation api

lightricks/ltx-2.0-pro/audio-to-video

Ltx 2.0 Pro Audio To Video by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Input
Text description of how the video should be generated. Required if image_url is not provided. When image_url is provided, this describes how the image should be animated.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of an image to use as the first frame of the video. If not provided, prompt is required.
URL of the audio file to generate a video from. Duration must be between 2 and 20 seconds. Must be publicly accessible or base64 data URI.
Idle

Example output — click Run to generate your own

API README

LTX 2.0 Pro Audio to Video

LTX 2.0 Pro Audio to Video is a audio-guided video creation endpoint in the LTX 2.0 family. It is built for creators who need to turn a concrete creative brief into a controlled visual sequence: the request establishes the source material, the intended subject behavior, the camera language, and the atmosphere of the finished shot. The production route keeps that workflow explicit instead of hiding its input assumptions behind a generic video-generation label.

In practice, this route accepts image, audio as creative context and exposes focused generation controls for delivery planning. That makes it suitable for shot-based pipelines where teams must preserve a source, direct a transformation, or control the final format without losing sight of the model's central task. Write prompts as a compact shot plan—subject, action, setting, camera, light, and timing—then use the structured fields for constraints that should remain deterministic across iterations.

Highlights

  • Audio-led timing. Uses the supplied soundtrack as a temporal signal while the prompt and opening image establish the visible scene.
  • Synchronized audio path. Audio conditioning or generation is available in the same request, helping sound and picture share one creative plan.
  • Temporal coherence. The LTX 2.0 generation path is designed around continuous motion across a shot, not a collection of unrelated frames.
  • Production iteration. Seed and sampling controls support repeatable comparisons when refining motion, framing, and visual treatment.

Pricing

The request price is calculated with params.duration * 0.10.

ConfigurationPrice
Formula-priced requestparams.duration * 0.10

The model card records a base price of $0.100000; the formula above determines usage-priced requests.

When to Use

ScenarioRecommendation
Choose this routeUse it when the deliverable specifically calls for audio-guided video creation, rather than a neighboring generation mode.
Prepare the sourceProvide image, audio in the format described by the request schema.
Direct the shotDescribe the subject, action, environment, camera movement, lighting, and temporal progression in that order.
Control continuityUse endpoint frames, reference media, strength, or audio controls when those fields are available instead of burying hard constraints in prose.
Plan deliverySet duration, frame count, resolution, and aspect ratio explicitly when the schema exposes them, then compare iterations with a stable seed where supported.

Prompt Guide

For audio-guided video creation, describe one coherent shot rather than a list of visual keywords. Put the main subject and action first, follow with location and staging, then add camera movement, lens or framing, lighting, mood, and any timed change. Keep URLs and hard delivery choices in their dedicated fields.

{
  "prompt": "An angry man speaking",
  "image": "https://static.sandbase.ai/examples/lightricks/ltx-2.0-pro/audio-to-video/input_image_0.png",
  "audio": "https://static.sandbase.ai/examples/lightricks/ltx-2.0-pro/audio-to-video/output_url_1.mp3"
}

Technical Specs

SpecificationValue
Model IDlightricks/ltx-2.0-pro/audio-to-video
Required inputsprompt, image
ExecutionAsynchronous generation job
Request controls3 documented fields
Outputurl, content_type

Request fields

FieldType and constraints
audiostring; Optional
imagestring; Required
promptstring; Required

Related Models

Related Models

lightricks/ltx-2.0-pro/extendLtx 2.0 Pro Extend by Lightricks - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.lightricks/ltx-2.0-pro/image-to-videoLtx 2.0 Pro by Lightricks - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.lightricks/ltx-2.0-pro/retakeLtx 2.0 Pro Retake by Lightricks - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.lightricks/ltx-2.0-pro/text-to-videoLtx 2.0 Pro is Lightricks's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.lightricks/ltx-2-19b/distilled/extend-video/loraLtx 2 19b Distilled Extend Video is Lightricks's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.lightricks/ltx-2-19b/distilled/image-to-videoLtx 2 19b Distilled by Lightricks - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.lightricks/ltx-2-19b/distilled/image-to-video/loraLtx 2 19b Distilled Lora is Lightricks's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.lightricks/ltx-2-19b/distilled/text-to-videoLtx 2 19b Distilled is Lightricks's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.