elevenlabs/text-to-dialogue
Text To Dialogue by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runelevenlabs/text-to-dialogueInput Schema
3 parameters · 0 required · 3 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
inputs | array | Optional | A list of dialogue inputs, each containing text and a voice ID which will be converted into speech. |
stability | number | Optional | Determines how stable the voice is and the randomness between each generation. Lower values introduce broader emotional range for the voice. Higher values can result in a monotonous voice with limited emotion. Must be one of 0.0, 0.5, 1.0, else it will be rounded to the nearest value. · Min: 0 · Max: 1 |
language_code | string | Optional | Language code (ISO 639-3) used to enforce a language for the model. · Options: , afr, ara, hye, asm, aze, bel, ben, bos, bul, cat, ceb, nya, hrv, ces, dan, nld, eng, est, fil, fin, fra, glg, kat, deu, ell, guj, hau, heb, hin, hun, isl, ind, gle, ita, jpn, jav, kan, kaz, kir, kor, lav, lin, lit, ltz, mkd, msa, mal, cmn, mar, nep, nor, pus, fas, pol, por, pan, ron, rus, srp, snd, slk, slv, som, spa, swa, swe, tam, tel, tha, tur, ukr, urd, vie, cym afrarahyeasmazebelbenbosbulcatcebnyahrvcesdannldengestfilfinfraglgkatdeuellgujhauhebhinhunislindgleitajpnjavkankazkirkorlavlinlitltzmkdmsamalcmnmarnepnorpusfaspolporpanronrussrpsndslkslvsomspaswaswetamtelthaturukrurdviecym |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "elevenlabs/text-to-dialogue",
"inputs": [
{
"text": "[applause] Thank you all for coming tonight! Today we have a very special guest with us.",
"voice": "Aria"
},
{
"text": "[gulps] ... [strong canadian accent] [excited] Hello everyone! Thank you all for having me tonight on this special day.",
"voice": "Charlotte"
}
],
"prompt": "a beautiful sunset over mountains"
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Text to Dialogue
elevenlabs/text-to-dialogue renders a scripted exchange as one coherent audio scene, allowing individual dialogue turns to use different voices while preserving conversational pacing. Rather than synthesizing isolated lines and manually assembling them, the model considers the sequence as a performance, producing more natural turn-taking, reactions, emphasis, and continuity between speakers. This combination makes the model a practical choice when the creative outcome depends on those qualities rather than on a generic media conversion.
For production work, It is useful for dramatized content, character prototypes, training scenarios, podcasts, and conversational ads where several distinct voices must sound as though they share the same scene. The result is most reliable when the source material and creative brief clearly describe the intended subject, progression, visual or sonic character, and the qualities that must remain unchanged.
Highlights
Multi-speaker performance assigns distinct voices to individual dialogue turns.
Conversational timing creates more natural pauses, interruptions, and turn-taking.
Scene-level continuity keeps energy and delivery coherent across the entire exchange.
Expressive dialogue rendering supports character-driven and dramatized audio.
Pricing
| Billing unit | Price |
|---|---|
Per dialogue turn in inputs | $0.100 |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| The model's named workflow matches the source material and intended output | A different input modality or model route is required |
| A managed asynchronous result is suitable for the production pipeline | A synchronous, interactive editor is essential |
| The documented controls cover the required duration, framing, or format | The project needs controls outside this endpoint's schema |
| Creative iteration benefits from a repeatable request structure | Exact deterministic pixels, frames, geometry, or samples are mandatory |
| A finished downloadable media asset is the desired deliverable | Editable source layers or a native project file are required |
Prompt Guide
For generation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.
{}
Technical Specs
| Spec | Value |
|---|---|
| Model ID | elevenlabs/text-to-dialogue |
| Inputs | inputs, language_code, stability |
| Required inputs | None |
| Output fields | content_type, url |
| Execution | Async (submit, then poll for result) |

