Audio Generation ​
SandBase currently publishes API reference pages for 46 enabled audio generation models across 10 providers. Choose a provider in the left navigation, then open a model page for its exact API identifier, supported capabilities, and a working request.
Audio Generation models use the async SandBase generation protocol declared in each model registry file. Submit a request, receive a task id, then poll the result endpoint until the generation is completed, failed, or timed out.
Providers ​
OpenAI ​
- Wizper (Whisper v3) — Wizper by OpenAI - accurate speech-to-text transcription with AI. Convert audio and video to text with high accuracy, multilingual support, and speaker identification.
ElevenLabs ​
- Scribe V2 — Scribe V2 is ElevenLabs's speech recognition model. Transcribe audio content with industry-leading accuracy across multiple languages and accents.
- Music — Music is ElevenLabs's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.
- Text To Dialogue — Text To Dialogue by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
- Sound Effects V2 — Sound Effects V2 by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
- V3 — V3 by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
- Turbo V2.5 — Turbo V2.5 by ElevenLabs - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
- Multilingual V2 — Multilingual V2 is ElevenLabs's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.
- ElevenLabs Audio Isolation — Audio Isolation by ElevenLabs - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- Voice Changer — Voice Changer by ElevenLabs - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
Google ​
- Gemini 3.1 Flash Tts — Gemini 3.1 Flash Tts by Google - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
- Gemini TTS — Gemini TTS by Google - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
- Lyria2 — Lyria 2 is Google's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.
MiniMax ​
- MiniMax Music 2.5 — MiniMax Music 2.5 text-to-music model with lyrics support, instrumental generation, and configurable audio output.
- MiniMax Music 2.6 — MiniMax Music 2.6 text-to-music model with enhanced quality, lyrics support, instrumental generation, and configurable audio output.
- MiniMax Speech 2.8 Turbo — MiniMax Speech 2.8 Turbo text-to-speech model with fast voice synthesis, enhanced expressiveness, 40+ language support, and voice cloning.
- Minimax Music — Music V2 by MiniMax - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
- MiniMax Speech 2.6 Turbo — MiniMax Speech 2.6 Turbo text-to-speech model with fast voice synthesis, 40+ language support, voice cloning, and expressive emotion control.
- MiniMax Speech 2.6 HD — MiniMax Speech 2.6 HD text-to-speech model with high-definition voice synthesis, 40+ language support, voice cloning, and expressive emotion control.
- MiniMax (Hailuo AI) Music v1.5 — Music V1.5 by MiniMax - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
- Minimax — Preview Speech 2.5 Hd is MiniMax's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
- MiniMax Speech 2.5 Turbo — Preview Speech 2.5 Turbo by MiniMax - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
- MiniMax Voice Design — Voice Design by MiniMax - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
- MiniMax Voice Cloning — Voice Clone is MiniMax's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
- MiniMax Speech-02 Turbo — Speech 02 Turbo is MiniMax's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
- …and 3 more models in the sidebar.
Mirelo ​
- Mirelo SFX1.6 — Sfx1.6 by Mirelo - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
stability-ai ​
- Stable Audio 2.5 Audio to Audio — Stable Audio 2.5 Audio To Audio by stability-ai - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- Stable Audio 2.5 Text to Audio — Stable Audio 2.5 by stability-ai - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
- Stable Audio 2.5 Inpaint — Stable Audio 2.5 Inpaint by stability-ai - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- Stable Audio Open — Stable Audio Open is stability-ai's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.
Alibaba ​
- Qwen 3 TTS - Voice Design [1.7B] — Qwen 3 Tts Voice Design 1.7b by Alibaba - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
- Qwen 3 TTS - Text to Speech [1.7B] — Qwen 3 Tts 1.7b is Alibaba's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
- Qwen 3 TTS - Text to Speech [0.6B] — Qwen 3 Tts 0.6b is Alibaba's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
- Qwen 3 TTS - Clone Voice [1.7B] — Qwen 3 Tts Clone Voice 1.7b by Alibaba - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- Qwen 3 TTS - Clone Voice [0.6B] — Qwen 3 Tts Clone Voice 0.6b by Alibaba - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
ace ​
- ACE-Step Audio Outpaint — Ace Step Audio Outpaint by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- ACE-Step — Ace Step Audio Inpaint by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- ACE-Step — Ace Step Audio To Audio by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- ACE-Step — Ace Step Prompt To Audio by ace - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
- ACE-Step — Ace Step by ace - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
KwaiVGI ​
- Kling Video — Kling Video Video To Audio by KwaiVGI - advanced AI model for video-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
- Kling TTS — Kling Tts is KwaiVGI's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
xAI ​
- xAI Text to Speech — Grok Tts by xAI - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
Capability coverage ​
audio-to-audio, speech-to-text, text-to-audio, text-to-music, text-to-speech, video-to-audio

