SandBase is live — $1 in free credits on signupStart free ›

Lightricks modelsvideo generation api

lightricks/ltx-2.3-22b/distilled/audio-to-video/lora

Ltx 2.3 22b Distilled Audio To Video by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Input
The prompt to generate the video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

Optional URL of an image to use as the first frame of the video.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the image to use as the end of the video.
The URL of the audio to generate the video from.
The output type of the generated video. Allowed values: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif).
160
The frames per second of the generated video. Range: 1 to 60.
01
The scale of the distill LoRA to use for the second and subsequent passes. Range: 0 to 1.
When enabled, the number of frames will be calculated based on the audio duration and FPS. When disabled, use the specified num_frames.
The quality of the generated video. Allowed values: low, medium, high, maximum.
The size of the generated video. Use 'auto' to match the input image dimensions if provided. Allowed values: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9.
The LoRAs to use for the generation.
01
The strength of the image to use for the video generation. Range: 0 to 1.
9481
The number of frames to generate. Range: 9 to 481.
01
The scale of the camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Range: 0 to 1.
01
Audio conditioning strength. Values below 1.0 will allow the model to change the audio, while a value of exactly 1.0 will use the input audio without modification. Range: 0 to 1.
01
The strength of the end image to use for the video generation. Range: 0 to 1.
The write mode of the generated video. Allowed values: fast, balanced, small.
Whether to use multi-scale generation. If True, the model will generate the video at a smaller scale first, then use the smaller video to guide the generation of a video at or above your requested size. This results in better coherence and details.
The scheduler to use. Allowed values: ltx2, linear_quadratic, beta.
The camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Allowed values: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none.
Whether to preprocess the audio before using it as conditioning.
The seed for the random number generator.
Idle

Example output — click Run to generate your own

API README

LTX-2.3 22B Distilled

LTX-2.3 22B Distilled is a audio-guided video creation with custom LoRA adapters endpoint in the LTX 2.3 family. It is built for creators who need to turn a concrete creative brief into a controlled visual sequence: the request establishes the source material, the intended subject behavior, the camera language, and the atmosphere of the finished shot. The distilled, LoRA-enabled route keeps that workflow explicit instead of hiding its input assumptions behind a generic video-generation label.

In practice, this route accepts image, audio, end image, loras as creative context and exposes num frames, video size, image strength, camera lora for delivery planning. That makes it suitable for shot-based pipelines where teams must preserve a source, direct a transformation, or control the final format without losing sight of the model's central task. Write prompts as a compact shot plan—subject, action, setting, camera, light, and timing—then use the structured fields for constraints that should remain deterministic across iterations.

Highlights

  • Soundtrack-led performance. The audio input supplies timing and performance cues, while the prompt explains what should happen visually.
  • Anchored opening composition. A required opening image establishes the subject and first-frame composition before motion follows the soundtrack.
  • Audio influence control. audio_strength is available on this route where documented, letting a request tune adherence to the supplied sound rather than changing the prompt.
  • Route-specific LoRA conditioning. The loras array adds learned styling or concepts to this audio to video workflow without replacing its required source inputs. For lightricks/ltx-2.3-22b/distilled/audio-to-video/lora, this task-specific control is documented alongside the route’s own required inputs.

Pricing

Billing unitPrice
Base request$0.001405

When to Use

ScenarioRecommendation
Choose this routeUse it when the deliverable specifically calls for audio-guided video creation with custom LoRA adapters, rather than a neighboring generation mode.
Prepare the sourceProvide image, audio in the format described by the request schema.
Direct the shotDescribe the subject, action, environment, camera movement, lighting, and temporal progression in that order.
Control continuityUse endpoint frames, reference media, strength, or audio controls when those fields are available instead of burying hard constraints in prose.
Plan deliverySet duration, frame count, resolution, and aspect ratio explicitly when the schema exposes them, then compare iterations with a stable seed where supported.

Prompt Guide

For audio-guided video creation with custom LoRA adapters, describe one coherent shot rather than a list of visual keywords. Put the main subject and action first, follow with location and staging, then add camera movement, lens or framing, lighting, mood, and any timed change. Keep URLs and hard delivery choices in their dedicated fields.

{
  "prompt": "A woman speaks to the camera",
  "image": "https://static.sandbase.ai/examples/lightricks/ltx-2.3-22b/distilled/audio-to-video/lora/input_image_1.png",
  "audio": "https://static.sandbase.ai/examples/lightricks/ltx-2.3-22b/distilled/audio-to-video/lora/input_audio_0.mp3",
  "end_image": "https://example.com/reference.jpg"
}

Technical Specs

SpecificationValue
Model IDlightricks/ltx-2.3-22b/distilled/audio-to-video/lora
Required inputsprompt, image
ExecutionAsynchronous generation job
Request controls22 documented fields
Outputurl, content_type

Request fields

FieldType and constraints
fpsnumber; Optional; Range: 1–60; Default: 24
seedinteger; Optional
audiostring; Optional
imagestring; Required
lorasarray; Optional
promptstring; Required
end_imagestring; Optional
schedulerstring; Optional; Options: ltx2, linear_quadratic, beta; Default: "ltx2"
num_framesinteger; Optional; Range: 9–481; Default: 121
video_sizestring; Optional; Options: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9; Default: "landscape_16_9"
camera_lorastring; Optional; Options: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none; Default: "none"
video_qualitystring; Optional; Options: low, medium, high, maximum; Default: "high"
audio_strengthnumber; Optional; Range: 0–1; Default: 1
image_strengthnumber; Optional; Range: 0–1; Default: 1
use_multiscaleboolean; Optional; Default: true
preprocess_audioboolean; Optional; Default: true
video_write_modestring; Optional; Options: fast, balanced, small; Default: "balanced"
camera_lora_scalenumber; Optional; Range: 0–1; Default: 1
video_output_typestring; Optional; Options: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif); Default: "X264 (.mp4)"
end_image_strengthnumber; Optional; Range: 0–1; Default: 1
match_audio_lengthboolean; Optional; Default: true
distill_lora_second_pass_scalenumber; Optional; Range: 0–1; Default: 0.5

Related Models

Related Models

lightricks/ltx-2.3-22b/distilled/image-to-video/loraLtx 2.3 22b Distilled Lora is Lightricks's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.lightricks/ltx-2.3-22b/distilled/reference-video-to-video/loraLtx 2.3 22b Distilled Reference Video To Video is Lightricks's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.lightricks/ltx-2.3-22b/distilled/text-to-video/loraLtx 2.3 22b Distilled Lora by Lightricks - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.lightricks/ltx-2.3-22b/distilled/video-to-video/loraLtx 2.3 22b Distilled Video To Video is Lightricks's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.lightricks/ltx-2.3-22b/audio-to-videoLtx 2.3 22b Audio To Video by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.lightricks/ltx-2.3-22b/audio-to-video/loraLtx 2.3 22b Audio To Video Lora by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.lightricks/ltx-2.3-22b/extendLtx 2.3 22b Extend by Lightricks - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.lightricks/ltx-2.3-22b/extend-video/loraLtx 2.3 22b Extend Video Lora is Lightricks's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.