Bytedance modelsaudio generation api

bytedance/seed-speech/tts/2.0

Seed Speech Tts 2.0 is Bytedance's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.

Input
The text to synthesize into speech.
Output audio format. 'mp3' returns MP3 audio; 'opus' returns Opus audio in an Ogg container. Allowed values: mp3, opus.
Sample rate of the output audio in Hz. Allowed values: 8000, 16000, 22050, 24000, 32000, 44100, 48000.
0.52
Volume. 1.0 is normal volume, 0.5 is half, 2.0 is double. Range: 0.5 to 2.
Voice to use for synthesis. The preset name encodes the voice and its supported language codes. 'mixed_en_zh' means the voice can seamlessly blend English and Chinese; separate codes (e.g. 'en_zh') mean the voice supports each language independently. Allowed values: vivi_mixed_en_zh_ja_es_id, mindy_en_es_id_pt_zh, stokie_en, dacey_en, tim_en, kian_en_zh, cedric_en_zh, sophie_en_zh, jean_en_zh, magnus_en_zh, mabel_en_zh, nadia_en_zh, opal_en_zh, pearl_en_zh, quentin_en_zh, vienna_mixed_en_zh, alina_mixed_en_zh, corinne_mixed_en_zh, esther_mixed_en_zh, freya_mixed_en_zh, gigi_mixed_en_zh, holly_mixed_en_zh, lyla_mixed_en_zh, daisy_mixed_en_zh, tracy_es_zh, jess_ja_es_id_pt_en_zh, pinky_es_ko_mixed_en_zh, sweety_ja_es, sandy_es_mixed_en_zh, sven_de, minimi_ja, usseau_fr, felipe_es, han_id, martins_pt, enzo_it, shane_ko, bonnie_zh, felix_zh, celeste_zh, monkey_king_zh.
-1212
Voice pitch shift in semitones. 0 is normal pitch, -12 lowers by one octave, 12 raises by one octave. Range: -12 to 12.
Force the text to be read as a single language, disabling automatic language detection. Leave unset for automatic detection (including seamless Chinese/English mixing on bilingual voices). Codes: zh (Chinese), en (English), ja (Japanese), es-mx (Mexican Spanish), id (Indonesian), pt-br (Brazilian Portuguese), ko (Korean), it (Italian), de (German), fr (French). Allowed values: zh, en, ja, es-mx, id, pt-br, ko, it, de, fr.
Optional natural-language instruction that steers the delivery (tone, emotion, pace, volume), e.g. 'Speak in a cheerful tone' or 'Could you speak a bit slower?'. It is not spoken aloud and does not affect billing.
0.52
Speech speed. 1.0 is normal speed, 0.5 is half speed, 2.0 is double speed. Range: 0.5 to 2.
Idle

Example output — click Run to generate your own

API README

PRICED DRAFT — MANUAL REVIEW REQUIRED exact_model_mode: bytedance/seed-speech/tts/2.0 original_unit: method: per-started-1000-characters price result_base_price: 0.03 result_price_formula: Math.ceil(len(params.text) / 1000) * 0.03 review_status: pending; model remains disabled checked_at: 2026-08-16

More Models by Bytedance

bytedance/seedance/2.0/image-to-videoByteDance's most advanced image-to-video model transforming still images into cinematic video with native audio, multi-shot editing, and director-level camera control for professional-grade video cre...Pay per usebytedance/seedream/4.5/editA new-generation image creation model from ByteDance, Seedream 4.5 integrates text-to-image generation and image editing into a single unified architecture, delivering high-fidelity visuals, precise p...Pay per usebytedance/seedance/2.0/reference-to-videoByteDance's most advanced reference-to-video model generating cinematic video guided by reference content, with native audio, multi-shot editing, and director-level camera control for professional-gr...Pay per usebytedance/seedvr/upscale/imageSeedvr Upscale Image by Bytedance - AI-powered image editing, style transfer, and transformation. Edit photos with natural language instructions, remove backgrounds, change styles, and enhance images ...Pay per usebytedance/seedream/5.0/lite/editSeedream 5.0 Lite — Fast Text-to-Image API The lightweight version of Seedream 5.0, delivering high-quality, low-latency AI image generation from text prompts. Ideal for real-time creative tools, e-co...Pay per usebytedance/seedance/2.0/text-to-videoByteDance's most advanced text-to-video model delivering cinematic output, native audio, multi-shot editing, and director-level camera control for professional-grade video creation.Pay per usebytedance/seedream/4.5A new-generation image creation model from ByteDance, Seedream 4.5 integrates text-to-image generation and image editing into a single unified architecture, delivering high-fidelity visuals, precise p...Pay per usebytedance/seedance/2.0/fast/image-to-videoByteDance's most advanced image-to-video model in its fast tier delivering lower latency and cost without compromising on cinematic output, native audio, multi-shot editing, and director-level camera...Pay per use