SandBase is live — $1 in free credits on signupStart free ›

Lightricks modelsimage generation api

lightricks/ltx-2-19b/distilled/audio-to-video

Ltx 2 19b Distilled Audio To Video by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Input
The prompt to generate the video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

Optional URL of an image to use as the first frame of the video.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the image to use as the end of the video.
The URL of the audio to generate the video from.
The quality of the generated video. Allowed values: low, medium, high, maximum.
The output type of the generated video. Allowed values: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif).
01
Audio conditioning strength. Values below 1.0 will allow the model to change the audio, while a value of exactly 1.0 will use the input audio without modification. Range: 0 to 1.
160
The frames per second of the generated video. Range: 1 to 60.
The size of the generated video. Use 'auto' to match the input image dimensions if provided. Allowed values: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9.
The write mode of the generated video. Allowed values: fast, balanced, small.
Whether to use multi-scale generation. If True, the model will generate the video at a smaller scale first, then use the smaller video to guide the generation of a video at or above your requested size. This results in better coherence and details.
01
The scale of the camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Range: 0 to 1.
01
The strength of the end image to use for the video generation. Range: 0 to 1.
The camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Allowed values: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none.
9481
The number of frames to generate. Range: 9 to 481.
When enabled, the number of frames will be calculated based on the audio duration and FPS. When disabled, use the specified num_frames.
01
The strength of the image to use for the video generation. Range: 0 to 1.
Whether to preprocess the audio before using it as conditioning.
The seed for the random number generator.
Idle

Example output — click Run to generate your own

API README

LTX-2 19B Distilled

LTX-2 19B Distilled is a audio-guided video creation endpoint in the LTX 2 19B family. It is built for creators who need to turn a concrete creative brief into a controlled visual sequence: the request establishes the source material, the intended subject behavior, the camera language, and the atmosphere of the finished shot. The distilled route keeps that workflow explicit instead of hiding its input assumptions behind a generic video-generation label.

In practice, this route accepts image, audio, end image as creative context and exposes num frames, video size, image strength, camera lora for delivery planning. That makes it suitable for shot-based pipelines where teams must preserve a source, direct a transformation, or control the final format without losing sight of the model's central task. Write prompts as a compact shot plan—subject, action, setting, camera, light, and timing—then use the structured fields for constraints that should remain deterministic across iterations.

Highlights

  • Soundtrack-led performance. The audio input supplies timing and performance cues, while the prompt explains what should happen visually.
  • Anchored opening composition. A required opening image establishes the subject and first-frame composition before motion follows the soundtrack.
  • Audio influence control. audio_strength is available on this route where documented, letting a request tune adherence to the supplied sound rather than changing the prompt.
  • Distilled-route planned final frame. An optional end_image can define the visual destination of the audio-driven shot. This capability is exposed by the named distilled route and its exact request fields. For lightricks/ltx-2-19b/distilled/audio-to-video, this task-specific control is documented alongside the route’s own required inputs.

Pricing

Billing unitPrice
Base request$0.000800

When to Use

ScenarioRecommendation
Choose this routeUse it when the deliverable specifically calls for audio-guided video creation, rather than a neighboring generation mode.
Prepare the sourceProvide image, audio in the format described by the request schema.
Direct the shotDescribe the subject, action, environment, camera movement, lighting, and temporal progression in that order.
Control continuityUse endpoint frames, reference media, strength, or audio controls when those fields are available instead of burying hard constraints in prose.
Plan deliverySet duration, frame count, resolution, and aspect ratio explicitly when the schema exposes them, then compare iterations with a stable seed where supported.

Prompt Guide

For audio-guided video creation, describe one coherent shot rather than a list of visual keywords. Put the main subject and action first, follow with location and staging, then add camera movement, lens or framing, lighting, mood, and any timed change. Keep URLs and hard delivery choices in their dedicated fields.

{
  "prompt": "A woman speaks to the camera",
  "image": "https://static.sandbase.ai/examples/lightricks/ltx-2-19b/distilled/audio-to-video/input_image_0.png",
  "audio": "https://static.sandbase.ai/examples/lightricks/ltx-2-19b/distilled/audio-to-video/input_audio_1.mp3",
  "end_image": "https://example.com/reference.jpg"
}

Technical Specs

SpecificationValue
Model IDlightricks/ltx-2-19b/distilled/audio-to-video
Required inputsprompt, image
ExecutionAsynchronous generation job
Request controls19 documented fields
Outputurl, content_type

Request fields

FieldType and constraints
fpsnumber; Optional; Range: 1–60; Default: 25
seedinteger; Optional
audiostring; Optional
imagestring; Required
promptstring; Required
end_imagestring; Optional
num_framesinteger; Optional; Range: 9–481; Default: 121
video_sizestring; Optional; Options: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9; Default: "landscape_4_3"
camera_lorastring; Optional; Options: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none; Default: "none"
video_qualitystring; Optional; Options: low, medium, high, maximum; Default: "high"
audio_strengthnumber; Optional; Range: 0–1; Default: 1
image_strengthnumber; Optional; Range: 0–1; Default: 1
use_multiscaleboolean; Optional; Default: true
preprocess_audioboolean; Optional; Default: true
video_write_modestring; Optional; Options: fast, balanced, small; Default: "balanced"
camera_lora_scalenumber; Optional; Range: 0–1; Default: 1
video_output_typestring; Optional; Options: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif); Default: "X264 (.mp4)"
end_image_strengthnumber; Optional; Range: 0–1; Default: 1
match_audio_lengthboolean; Optional; Default: true

Related Models

Related Models