SandBase is live — $1 in free credits on signupStart free ›

Lightricks modelsvideo generation api

lightricks/ltx-2.3-22b/audio-to-video/lora

Ltx 2.3 22b Audio To Video Lora by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Input
The prompt to generate the video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

Optional URL of an image to use as the first frame of the video.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

The URL of the image to use as the end of the video.
The URL of the audio to generate the video from.
850
The number of inference steps to use. Range: 8 to 50.
When enabled, the number of frames will be calculated based on the audio duration and FPS. When disabled, use the specified num_frames.
01
The strength of the image to use for the video generation. Range: 0 to 1.
160
The frames per second of the generated video. Range: 1 to 60.
The quality of the generated video. Allowed values: low, medium, high, maximum.
01
The rescaling scale for the video. Controls the ratio between classifier-free guidance and spatiotemporal guidance. Range: 0 to 1.
The camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Allowed values: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none.
The scheduler to use. Allowed values: ltx2, linear_quadratic, beta.
010
The modality scale for the audio. Controls the ratio between video and audio modalities. Range: 0 to 10.
01
The scale of the distill LoRA to use for the first pass. Set to 0 to disable. Range: 0 to 1.
The write mode of the generated video. Allowed values: fast, balanced, small.
010
The gamma of gradient estimation during denoising. Set to 0 to disable. Range: 0 to 10.
010
The modality scale for the video. Controls the ratio between video and audio modalities. Range: 0 to 10.
01
Audio conditioning strength. Values below 1.0 will allow the model to change the audio, while a value of exactly 1.0 will use the input audio without modification. Range: 0 to 1.
The LoRAs to use for the generation.
The output type of the generated video. Allowed values: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif).
120
The Classifier-Free Guidance (CFG) scale for the video. Higher values result in more consistent and focused video content. Range: 1 to 20.
01
The rescaling scale for the audio. Controls the ratio between classifier-free guidance and spatiotemporal guidance. Range: 0 to 1.
Whether to use multi-scale generation. If True, the model will generate the video at a smaller scale first, then use the smaller video to guide the generation of a video at or above your requested size. This results in better coherence and details.
120
The Classifier-Free Guidance (CFG) scale for the audio. Higher values result in more consistent and focused audio content. Range: 1 to 20.
020
The Spatiotemporal Guidance (STG) scale for the audio. Higher values result in more consistent and focused audio content. Range: 0 to 20.
Whether to use restart sampling. This will inject a small amount of noise during each denoising step, which can help improve the quality of the generated video.
9481
The number of frames to generate. Range: 9 to 481.
Whether to preprocess the audio before using it as conditioning.
01
The scale of the distill LoRA to use for the second and subsequent passes. Range: 0 to 1.
The size of the generated video. Use 'auto' to match the input image dimensions if provided. Allowed values: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9.
020
The Spatiotemporal Guidance (STG) scale for the video. Higher values result in more consistent and focused video content. Range: 0 to 20.
01
The strength of the end image to use for the video generation. Range: 0 to 1.
01
The scale of the camera LoRA to use. This allows you to control the camera movement of the generated video more accurately than just prompting the model to move the camera. Range: 0 to 1.
The seed for the random number generator.
Idle

Example output — click Run to generate your own

API README

LTX-2.3 22B

LTX-2.3 22B is a audio-guided video creation with custom LoRA adapters endpoint in the LTX 2.3 family. It is built for creators who need to turn a concrete creative brief into a controlled visual sequence: the request establishes the source material, the intended subject behavior, the camera language, and the atmosphere of the finished shot. The LoRA-enabled route keeps that workflow explicit instead of hiding its input assumptions behind a generic video-generation label.

In practice, this route accepts image, audio, end image, loras as creative context and exposes num frames, video size, image strength, camera lora for delivery planning. That makes it suitable for shot-based pipelines where teams must preserve a source, direct a transformation, or control the final format without losing sight of the model's central task. Write prompts as a compact shot plan—subject, action, setting, camera, light, and timing—then use the structured fields for constraints that should remain deterministic across iterations.

Highlights

  • Soundtrack-led performance. The audio input supplies timing and performance cues, while the prompt explains what should happen visually.
  • Anchored opening composition. A required opening image establishes the subject and first-frame composition before motion follows the soundtrack.
  • Audio influence control. audio_strength is available on this route where documented, letting a request tune adherence to the supplied sound rather than changing the prompt.
  • Route-specific LoRA conditioning. The loras array adds learned styling or concepts to this audio to video workflow without replacing its required source inputs. For lightricks/ltx-2.3-22b/audio-to-video/lora, this task-specific control is documented alongside the route’s own required inputs.

Pricing

Billing unitPrice
Base request$0.001805

When to Use

ScenarioRecommendation
Choose this routeUse it when the deliverable specifically calls for audio-guided video creation with custom LoRA adapters, rather than a neighboring generation mode.
Prepare the sourceProvide image, audio in the format described by the request schema.
Direct the shotDescribe the subject, action, environment, camera movement, lighting, and temporal progression in that order.
Control continuityUse endpoint frames, reference media, strength, or audio controls when those fields are available instead of burying hard constraints in prose.
Plan deliverySet duration, frame count, resolution, and aspect ratio explicitly when the schema exposes them, then compare iterations with a stable seed where supported.

Prompt Guide

For audio-guided video creation with custom LoRA adapters, describe one coherent shot rather than a list of visual keywords. Put the main subject and action first, follow with location and staging, then add camera movement, lens or framing, lighting, mood, and any timed change. Keep URLs and hard delivery choices in their dedicated fields.

{
  "prompt": "A woman speaks to the camera",
  "image": "https://static.sandbase.ai/examples/lightricks/ltx-2.3-22b/audio-to-video/lora/input_image_1.png",
  "audio": "https://static.sandbase.ai/examples/lightricks/ltx-2.3-22b/audio-to-video/lora/input_audio_0.mp3",
  "end_image": "https://example.com/reference.jpg"
}

Technical Specs

SpecificationValue
Model IDlightricks/ltx-2.3-22b/audio-to-video/lora
Required inputsprompt, image
ExecutionAsynchronous generation job
Request controls34 documented fields
Outputurl, content_type

Request fields

FieldType and constraints
fpsnumber; Optional; Range: 1–60; Default: 24
seedinteger; Optional
audiostring; Optional
imagestring; Required
lorasarray; Optional
promptstring; Required
end_imagestring; Optional
schedulerstring; Optional; Options: ltx2, linear_quadratic, beta; Default: "ltx2"
num_framesinteger; Optional; Range: 9–481; Default: 121
video_sizestring; Optional; Options: auto, square_hd, square, portrait_4_3, portrait_16_9, landscape_4_3, landscape_16_9; Default: "landscape_16_9"
camera_lorastring; Optional; Options: dolly_in, dolly_out, dolly_left, dolly_right, jib_up, jib_down, static, none; Default: "none"
video_qualitystring; Optional; Options: low, medium, high, maximum; Default: "high"
audio_strengthnumber; Optional; Range: 0–1; Default: 1
image_strengthnumber; Optional; Range: 0–1; Default: 1
use_multiscaleboolean; Optional; Default: true
audio_cfg_scalenumber; Optional; Range: 1–20; Default: 7
audio_stg_scalenumber; Optional; Range: 0–20; Default: 0
video_cfg_scalenumber; Optional; Range: 1–20; Default: 3
video_stg_scalenumber; Optional; Range: 0–20; Default: 0
preprocess_audioboolean; Optional; Default: true
video_write_modestring; Optional; Options: fast, balanced, small; Default: "balanced"
camera_lora_scalenumber; Optional; Range: 0–1; Default: 1
video_output_typestring; Optional; Options: X264 (.mp4), VP9 (.webm), PRORES4444 (.mov), GIF (.gif); Default: "X264 (.mp4)"
end_image_strengthnumber; Optional; Range: 0–1; Default: 1
match_audio_lengthboolean; Optional; Default: true
num_inference_stepsinteger; Optional; Range: 8–50; Default: 40
audio_modality_scalenumber; Optional; Range: 0–10; Default: 3
use_restart_samplingboolean; Optional; Default: false
video_modality_scalenumber; Optional; Range: 0–10; Default: 3
audio_rescaling_scalenumber; Optional; Range: 0–1; Default: 0.7
video_rescaling_scalenumber; Optional; Range: 0–1; Default: 0.7
gradient_estimation_gammanumber; Optional; Range: 0–10; Default: 2
distill_lora_first_pass_scalenumber; Optional; Range: 0–1; Default: 0.2
distill_lora_second_pass_scalenumber; Optional; Range: 0–1; Default: 0.5

Related Models

Related Models

lightricks/ltx-2.3-22b/audio-to-videoLtx 2.3 22b Audio To Video by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.lightricks/ltx-2.3-22b/distilled/audio-to-video/loraLtx 2.3 22b Distilled Audio To Video by Lightricks - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.lightricks/ltx-2.3-22b/distilled/image-to-video/loraLtx 2.3 22b Distilled Lora is Lightricks's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.lightricks/ltx-2.3-22b/distilled/reference-video-to-video/loraLtx 2.3 22b Distilled Reference Video To Video is Lightricks's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.lightricks/ltx-2.3-22b/distilled/text-to-video/loraLtx 2.3 22b Distilled Lora by Lightricks - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.lightricks/ltx-2.3-22b/distilled/video-to-video/loraLtx 2.3 22b Distilled Video To Video is Lightricks's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.lightricks/ltx-2.3-22b/extendLtx 2.3 22b Extend by Lightricks - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.lightricks/ltx-2.3-22b/extend-video/loraLtx 2.3 22b Extend Video Lora is Lightricks's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.