xai/grok-tts
Grok Tts by xAI - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
Example output — click Run to generate your own
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
https://api.sandbase.ai/v1/runxai/grok-ttsInput Schema
2 parameters · 0 required · 2 optional
| Parameter | Type | Required | Description |
|---|---|---|---|
text | string | Optional | The text to convert to speech. Maximum 15,000 characters. Supports speech tags for expressive delivery: inline tags like [laugh], [pause], [sigh] and wrapping tags like <whisper>text</whisper>, <slow>text</slow>. · Min length: 1 · Max length: 15000 |
voice | string | Optional | The voice to use for speech synthesis. · Options: eve, ara, rex, sal, leo eveararexsalleo |
Output Schema
| Field | Type | Description |
|---|---|---|
id | string | Unique identifier for the generation task |
status | string | Task status: pending, running, completed, failed, timeout |
model | string | Model used for the generation |
outputs | array | Array of output items |
outputs[].url | string | URL of the generated artifact |
outputs[].content_type | string | MIME type (e.g. image/png, video/mp4) |
error | object | null | Error details if failed, null on success |
error.type | string | Machine-readable error type code |
error.message | string | Human-readable error description |
Async Workflow
This model uses asynchronous execution. Submit a request and poll for the result.
- Submit — POST to /v1/run, receive an
id - Poll — GET /v1/run/{id} until status is
completed,failed, ortimeout - Retrieve — Read
outputsfrom the completed response
Code Examples
Ready-to-run snippets
# Step 1: Submit
curl -X POST https://api.sandbase.ai/v1/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "xai/grok-tts",
"text": "Hello! This is xAI text to speech, brought to you by Fal AI.",
"prompt": "a beautiful sunset over mountains"
}'
# Step 2: Poll result (replace <id>)
curl https://api.sandbase.ai/v1/run/<id> \
-H "Authorization: Bearer YOUR_API_KEY"API README
Grok Text to Speech
Grok Text to Speech is a specialized SandBase endpoint built to turn scripts as long as 15,000 characters into expressive spoken audio. It is most useful when teams need five selectable voices, inline performance cues, and long-form narration in a single request, with the route’s inputs defining a repeatable production contract instead of leaving critical delivery choices to an ad-hoc manual workflow.
Use this exact route when its input mode matches the creative asset already in hand: describe the visual or spoken result clearly, set only the controls that support the intended delivery, and keep the subject, action, environment, camera or performance direction internally consistent. The result is returned asynchronously as a downloadable media URL suitable for review, automation, or downstream finishing.
Highlights
- Expressive speech markup. Inline cues such as
[laugh],[pause], and[sigh]add performance beats without splitting the script into separate clips. - Phrase-level delivery control. Wrapping tags such as
<whisper>and<slow>shape how selected passages are spoken while the surrounding narration remains natural. - Five distinct voices. Choose Eve, Ara, Rex, Sal, or Leo to match the tone and character of the production.
- Long-form synthesis. Inputs up to 15,000 characters support explainers, accessibility narration, lessons, and story passages in one request.
Pricing
Text is billed in 1,000-character blocks, rounded up. Each started block costs $0.0042.
| Text length | Price |
|---|---|
| 1–1,000 characters | $0.0042 |
| 1,001–2,000 characters | $0.0084 |
| 2,001–3,000 characters | $0.0126 |
| Up to 15,000 characters | $0.0630 |
When to Use
| Scenario | Why this route fits |
|---|---|
| Concept development | Choose Grok Text to Speech when its grok tts workflow matches the starting material and you need several clearly directed variations. |
| Production iteration | Use explicit duration, resolution, ratio, seed, or quality controls to compare versions without changing the core creative brief. |
| Channel adaptation | Generate directly in the landscape, square, portrait, or vertical format required by the destination whenever that control is available. |
| Automated pipelines | Integrate the asynchronous media URL into review queues, asset libraries, publishing tools, or a later finishing stage. |
| Alternative route | Pick a related text-, image-, reference-, edit-, or turbo route when the available source media or required degree of control is different. |
Prompt Guide
Lead with the main subject or source asset, then describe the intended action or transformation, environment, composition, camera or vocal delivery, lighting and mood. Keep instructions concrete and compatible; use the route’s explicit fields for duration, resolution, ratio, quality, voice, or reproducibility instead of burying those settings in prose.
{
"text": "Welcome to the product tour. [pause] Let us begin.",
"voice": "eve"
}
Technical Specs
| Property | Details |
|---|---|
| Model ID | xai/grok-tts |
| Required inputs | None declared |
| Execution | Asynchronous; poll the returned generation ID |
| Output | Downloadable media URL |
text | string; optional |
voice | string; optional; choices: eve, ara, rex, sal, leo |
Related Models
| Model | Best for |
|---|---|
xai/grok-imagine-video/text-to-video | Alternative text to video workflow |
xai/grok-imagine-video/extend | Alternative extend workflow |
xai/grok-imagine-video/image-to-video | Alternative image to video workflow |
xai/grok-imagine-video/reference-to-video | Alternative reference to video workflow |

