SandBase is live — $1 in free credits on signupStart free ›

ElevenLabs modelsvideo generation api

elevenlabs/dubbing

Dubbing by ElevenLabs - advanced AI model for audio-to-video. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Input
URL of the video file to dub. Either audio_url or video_url must be provided. If both are provided, video_url takes priority.
URL of the audio file to dub. Either audio_url or video_url must be provided.
Source language code. If not provided, will be auto-detected.
150
Number of speakers in the audio. If not provided, will be auto-detected. Range: 1 to 50.
Target language code for dubbing (ISO 639-1)
Whether to use the highest resolution for dubbing.
Idle

Example output — click Run to generate your own

API README

ElevenLabs Dubbing

ElevenLabs Dubbing localizes existing audio or video into another language while preserving the performance qualities that make each speaker recognizable. The dubbing workflow carries tone, pacing, delivery, and emotional intent across languages instead of producing a flat reading of a translated transcript.

It is designed for podcasts, interviews, creator videos, training material, and campaigns that need to reach new audiences without re-recording every participant. Multiple speakers can be handled within one source, while background music, effects, and ambience remain part of the localized result.

Highlights

  • Performance-preserving translation. Carries speaker tone, pacing, delivery, and emotion into the target-language performance.
  • Multiple-speaker handling. Separates and preserves distinct voices across conversations and overlapping dialogue.
  • Background-audio retention. Keeps music, ambience, and effects so the localized track does not require a full remix.
  • Audio and video input. Accepts existing media for localization rather than requiring a text-only narration workflow.

Pricing

ConfigurationBilling unitPrice
Source durationPer second$0.015

When to Use

✅ Good fit❌ Consider alternatives
The project needs localizing existing spoken mediaThe goal is a different media task or endpoint
The available inputs match the required local schemaRequired source media or permissions are unavailable
The brief can specify subject, composition, style, and deliveryThe result must be deterministic at pixel or sample level
The supported formats and controls match final placementDelivery requires unsupported dimensions, codecs, or duration
An asynchronous generated result fits the workflowA live, frame-synchronous, or real-time response is mandatory

Prompt Guide

Write the request as a production brief: identify the main subject or source material, state the intended transformation, describe composition or timing, and finish with style, atmosphere, and delivery constraints. Keep preservation requirements separate from requested changes, and use only fields exposed by this route.

{
  "audio": "https://example.com/reference.mp3",
  "video": "https://example.com/source.mp4",
  "source_lang": "example"
}

Technical Specs

SpecValue
Model IDelevenlabs/dubbing
Input fieldsaudio (string)<br>video (string)<br>source_lang (string)<br>target_lang (string)<br>num_speakers (integer; 1–50)<br>highest_resolution (boolean)
Required inputNone marked required
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

  • elevenlabs/sound-effects-v2

More Models by ElevenLabs

elevenlabs/v3V3 by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.Pay per useelevenlabs/sound-effects-v2Sound Effects V2 by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.Pay per useelevenlabs/turbo-v2.5Turbo V2.5 by ElevenLabs - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.Pay per useelevenlabs/multilingual-v2Multilingual V2 is ElevenLabs's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.Pay per useelevenlabs/musicMusic is ElevenLabs's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.Pay per useelevenlabs/scribe-v2Scribe V2 is ElevenLabs's speech recognition model. Transcribe audio content with industry-leading accuracy across multiple languages and accents.Pay per useelevenlabs/speech-to-textSpeech To Text by ElevenLabs - accurate speech-to-text transcription with AI. Convert audio and video to text with high accuracy, multilingual support, and speaker identification.Pay per useelevenlabs/audio-isolationAudio Isolation by ElevenLabs - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.Pay per use