SandBase is live — $1 in free credits on signupStart free ›

PixVerse modelsvideo generation api

pixverse/lipsync

Lipsync is PixVerse's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.

Input
URL of the input video
URL of the input audio. If not provided, TTS will be used.
Text content for TTS when audio_url is not provided
Voice to use for TTS when audio_url is not provided Allowed values: Emily, James, Isabella, Liam, Chloe, Adrian, Harper, Ava, Sophia, Julia, Mason, Jack, Oliver, Ethan, Auto.
Idle

Example output — click Run to generate your own

API README

PixVerse

PixVerse is the PixVerse route for lip synchronization, turning existing footage plus speech into a production-ready result while keeping the operation distinct from neighboring endpoints. Lipsync is PixVerse's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification. The workflow is designed for creators who need the model’s specific transformation to remain visible in the request, so the source, intended change, and finished artifact can be reviewed as one coherent creative decision.

In practical use, this route exposes text, voice_id, audio_url, sync_mode, video_url, safety_tolerance to shape the exact deliverable. Those controls let a team preserve the important source constraints, state subject behavior or material treatment precisely, choose supported timing or output characteristics, and reproduce successful settings across alternate takes. The result fits an iterative pipeline: establish the core brief, compare controlled variations, then pass the selected asset into editorial, design, localization, visualization, or publishing work.

Highlights

  • Audio-led mouth timing. Synchronize facial motion in an existing video to a finished spoken recording.
  • Built-in speech option. Supply text instead of audio and select a documented voice for a combined speech-and-lip-sync workflow.
  • Existing-footage workflow. Preserve the supplied performance and framing while revising the spoken delivery.
  • Dubbing-ready output. Create a synchronized clip that can move directly into review, localization, or social publishing.

Pricing

ConfigurationPrice
Video processing$0.040000 per second
Uploaded-audio workflowNo text-to-speech surcharge
Text-to-speech$0.240000 per started 100 characters

When to Use

ScenarioWhy it fits
Exact workflow fitChoose this route when the required deliverable is lip synchronization, rather than a related route with different source media.
Directed creative iterationUse it when subject, motion, material, speech, framing, or finish should be expressed explicitly and compared across controlled variants.
Existing-asset continuityUse it when supplied images, video, audio, references, or styles must remain the anchor for the generated result.
Repeatable productionUse it when successful inputs need to be saved and rerun across a campaign, asset set, localization pass, or batch.
Pipeline handoffUse it when the returned artifact will move into editorial, compositing, visualization, review, storage, or publishing.

Prompt Guide

Start with the source or subject, then describe the intended transformation, movement or behavior, camera and composition, and the desired finish. Keep media URLs reachable, use only fields documented for this exact route, and change one major control at a time when comparing results. For source-led tasks, describe what should change as well as what must remain recognizable.

{
  "text": "A cinematic close shot with clear subject action, camera movement, lighting, and final visual treatment."
}

Technical Specs

PropertyValue
Model IDpixverse/lipsync
Execution modeasync
Required inputsNone marked required
textstring
voice_idstring; options: Emily, James, Isabella, Liam, Chloe, Adrian, Harper, Ava, Sophia, Julia, Mason, Jack, Oliver, Ethan, Auto; default: auto
audio_urlstring
sync_modeboolean; default: false
video_urlstring
safety_tolerancestring; default: 6

Related Models

  • pixverse/extend
  • pixverse/sound-effects
  • pixverse/swap
  • pixverse/c1/image-to-video

Related Models

pixverse/c1/image-to-videoC1 is PixVerse's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.pixverse/c1/reference-to-videoC1 is PixVerse's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.pixverse/c1/text-to-videoC1 by PixVerse - generate cinematic videos from text descriptions with AI. Create high-quality video content with natural motion, camera control, and optional audio generation.pixverse/c1/transitionC1 Transition is PixVerse's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.pixverse/extendExtend by PixVerse - AI-powered video editing and transformation. Apply style transfer, motion control, lip-sync, and visual effects to existing videos with natural language instructions.pixverse/extend/fastExtend Fast is PixVerse's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.pixverse/sound-effectsSound Effects is PixVerse's video-to-video AI model. Transform, enhance, and edit video content using text prompts - from style changes to object manipulation and scene modification.pixverse/swapSwap by PixVerse - animate still images into dynamic videos with AI. Transform photos into cinematic clips with natural motion, camera movement, and optional audio generation.