SandBase is live — $1 in free credits on signupStart free ›

Lightricks modelsimage generation api

lightricks/ltx-2-19b/audio-to-video

Ltx 2 19b Audio To Video by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Input
The prompt to generate the video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

Optional URL of an image to use as the first frame of the video.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the image to use as the end of the video.
The URL of the audio to generate the video from.
110
The guidance scale to use. Range: 1 to 10.
850
The number of inference steps to use. Range: 8 to 50.
The size of the generated video. Use 'auto' to match the input image dimensions if provided. Allowed values: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9.
When enabled, the number of frames will be calculated based on the audio duration and FPS. When disabled, use the specified num_frames.
The output type of the generated video. Allowed values: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif).
01
Audio conditioning strength. Values below 1.0 will allow the model to change the audio, while a value of exactly 1.0 will use the input audio without modification. Range: 0 to 1.
01
The scale of the camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Range: 0 to 1.
The camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Allowed values: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none.
The write mode of the generated video. Allowed values: fast, balanced, small.
01
The strength of the end image to use for the video generation. Range: 0 to 1.
Whether to preprocess the audio before using it as conditioning.
160
The frames per second of the generated video. Range: 1 to 60.
9481
The number of frames to generate. Range: 9 to 481.
The quality of the generated video. Allowed values: low, medium, high, maximum.
Whether to use multi-scale generation. If True, the model will generate the video at a smaller scale first, then use the smaller video to guide the generation of a video at or above your requested size. This results in better coherence and details.
01
The strength of the image to use for the video generation. Range: 0 to 1.
The seed for the random number generator.
Idle

Example output — click Run to generate your own

API README

LTX-2 19B

LTX-2 19B is a audio-guided video creation endpoint in the LTX 2 19B family. It is built for creators who need to turn a concrete creative brief into a controlled visual sequence: the request establishes the source material, the intended subject behavior, the camera language, and the atmosphere of the finished shot. The standard route keeps that workflow explicit instead of hiding its input assumptions behind a generic video-generation label.

In practice, this route accepts image, audio, end image as creative context and exposes num frames, video size, image strength, camera lora for delivery planning. That makes it suitable for shot-based pipelines where teams must preserve a source, direct a transformation, or control the final format without losing sight of the model's central task. Write prompts as a compact shot plan—subject, action, setting, camera, light, and timing—then use the structured fields for constraints that should remain deterministic across iterations.

Highlights

  • Soundtrack-led performance. The audio input supplies timing and performance cues, while the prompt explains what should happen visually.
  • Anchored opening composition. A required opening image establishes the subject and first-frame composition before motion follows the soundtrack.
  • Audio influence control. audio_strength is available on this route where documented, letting a request tune adherence to the supplied sound rather than changing the prompt.
  • LTX 2 19B planned final frame. An optional end_image can define the visual destination of the audio-driven shot. This behavior is exposed on the named 19B route. For lightricks/ltx-2-19b/audio-to-video, this task-specific control is documented alongside the route’s own required inputs.

Pricing

Billing unitPrice
Base request$0.001800

When to Use

ScenarioRecommendation
Choose this routeUse it when the deliverable specifically calls for audio-guided video creation, rather than a neighboring generation mode.
Prepare the sourceProvide image, audio in the format described by the request schema.
Direct the shotDescribe the subject, action, environment, camera movement, lighting, and temporal progression in that order.
Control continuityUse endpoint frames, reference media, strength, or audio controls when those fields are available instead of burying hard constraints in prose.
Plan deliverySet duration, frame count, resolution, and aspect ratio explicitly when the schema exposes them, then compare iterations with a stable seed where supported.

Prompt Guide

For audio-guided video creation, describe one coherent shot rather than a list of visual keywords. Put the main subject and action first, follow with location and staging, then add camera movement, lens or framing, lighting, mood, and any timed change. Keep URLs and hard delivery choices in their dedicated fields.

{
  "prompt": "A woman speaks to the camera",
  "image": "https://static.sandbase.ai/examples/lightricks/ltx-2-19b/audio-to-video/input_image_0.png",
  "audio": "https://static.sandbase.ai/examples/lightricks/ltx-2-19b/audio-to-video/input_audio_1.mp3",
  "end_image": "https://example.com/reference.jpg"
}

Technical Specs

SpecificationValue
Model IDlightricks/ltx-2-19b/audio-to-video
Required inputsprompt, image
ExecutionAsynchronous generation job
Request controls21 documented fields
Outputurl, content_type

Request fields

FieldType and constraints
fpsnumber; Optional; Range: 1–60; Default: 25
seedinteger; Optional
audiostring; Optional
imagestring; Required
promptstring; Required
end_imagestring; Optional
num_framesinteger; Optional; Range: 9–481; Default: 121
video_sizestring; Optional; Options: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9; Default: "landscape_4_3"
camera_lorastring; Optional; Options: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none; Default: "none"
video_qualitystring; Optional; Options: low, medium, high, maximum; Default: "high"
audio_strengthnumber; Optional; Range: 0–1; Default: 1
guidance_scalenumber; Optional; Range: 1–10; Default: 3
image_strengthnumber; Optional; Range: 0–1; Default: 1
use_multiscaleboolean; Optional; Default: true
preprocess_audioboolean; Optional; Default: true
video_write_modestring; Optional; Options: fast, balanced, small; Default: "balanced"
camera_lora_scalenumber; Optional; Range: 0–1; Default: 1
video_output_typestring; Optional; Options: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif); Default: "X264 (.mp4)"
end_image_strengthnumber; Optional; Range: 0–1; Default: 1
match_audio_lengthboolean; Optional; Default: true
num_inference_stepsinteger; Optional; Range: 8–50; Default: 40

Related Models

Related Models