SandBase is live — $1 in free credits on signupStart free ›

OpenAI modelsaudio generation api

openai/wizper

Wizper by OpenAI - accurate speech-to-text transcription with AI. Convert audio and video to text with high accuracy, multilingual support, and speaker identification.

Input
URL of the audio file to transcribe. Supported formats: mp3, mp4, mpeg, mpga, m4a, wav or webm.
Task to perform on the audio file. Either transcribe or translate. Allowed values: transcribe, translate.
Level of the chunks to return.
1029
Maximum speech segment duration in seconds before splitting. Range: 10 to 29.
Language of the audio file. If translate is selected as the task, the audio will be translated to English, regardless of the language selected. If `None` is passed, the language will be automatically detected. This will also increase the inference time. Allowed values: af, am, ar, as, az, ba, be, bg, bn, bo, br, bs, ca, cs, cy, da, de, el, en, es, et, eu, fa, fi, fo, fr, gl, gu, ha, haw, he, hi, hr, ht, hu, hy, id, is, it, ja, jw, ka, kk, km, kn, ko, la, lb, ln, lo, lt, lv, mg, mi, mk, ml, mn, mr, ms, mt, my, ne, nl, nn, no, oc, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, sn, so, sq, sr, su, sv, sw, ta, te, tg, th, tk, tl, tr, tt, uk, ur, uz, vi, yi, yo, zh.
Version of the model to use. All of the models are the Whisper large variant.
Whether to merge consecutive chunks. When enabled, chunks are merged if their combined duration does not exceed max_segment_len.
Idle

Example output — click Run to generate your own

API README

Wizper

Wizper is the Wizper route for transcription, turning spoken audio or video into a production-ready result while keeping the operation distinct from neighboring endpoints. Wizper by OpenAI - accurate speech-to-text transcription with AI. Convert audio and video to text with high accuracy, multilingual support, and speaker identification. The workflow is designed for creators who need the model’s specific transformation to remain visible in the request, so the source, intended change, and finished artifact can be reviewed as one coherent creative decision.

In practical use, this route exposes task, version, language, audio_url, sync_mode, chunk_level, merge_chunks, max_segment_len to shape the exact deliverable. Those controls let a team preserve the important source constraints, state subject behavior or material treatment precisely, choose supported timing or output characteristics, and reproduce successful settings across alternate takes. The result fits an iterative pipeline: establish the core brief, compare controlled variations, then pass the selected asset into editorial, design, localization, visualization, or publishing work.

Highlights

  • Transcribe or translate. Turn supported audio or video into source-language text, or translate the spoken content into English.
  • Broad language coverage. Choose from the documented language codes while retaining explicit control over the recognition task.
  • Segment-aware processing. Tune chunk level, maximum segment length, and consecutive-chunk merging for long recordings.
  • Media-ready ingestion. Accept common MP3, MP4, MPEG, MPGA, M4A, WAV, and WebM assets through one transcription route.

Pricing

ConfigurationPrice
Per request$0.006000

When to Use

ScenarioWhy it fits
Exact workflow fitChoose this route when the required deliverable is transcription, rather than a related route with different source media.
Directed creative iterationUse it when subject, motion, material, speech, framing, or finish should be expressed explicitly and compared across controlled variants.
Existing-asset continuityUse it when supplied images, video, audio, references, or styles must remain the anchor for the generated result.
Repeatable productionUse it when successful inputs need to be saved and rerun across a campaign, asset set, localization pass, or batch.
Pipeline handoffUse it when the returned artifact will move into editorial, compositing, visualization, review, storage, or publishing.

Prompt Guide

Start with the source or subject, then describe the intended transformation, movement or behavior, camera and composition, and the desired finish. Keep media URLs reachable, use only fields documented for this exact route, and change one major control at a time when comparing results. For source-led tasks, describe what should change as well as what must remain recognizable.

{}

Technical Specs

PropertyValue
Model IDopenai/wizper
Execution modeasync
Required inputsNone marked required
taskstring; options: transcribe, translate; default: transcribe
versionstring; default: 3
languagestring; default: en
audio_urlstring
sync_modeboolean; default: false
chunk_levelstring; default: segment
merge_chunksboolean; default: true
max_segment_leninteger; minimum: 10; maximum: 29; default: 29

Related Models

  • openai/gpt-image-1
  • openai/gpt-image-1-mini
  • openai/gpt-image-1.5
  • openai/gpt-image-2

More Models by OpenAI

openai/gpt-image-2/editGPT Image 2 Editing supports image editing and multi-image synthesis with high-quality results.Pay per useopenai/gpt-image-2GPT Image 2, OpenAI's latest image model, is capable of making fine-grained, detailed edits to images.Pay per useopenai/gpt-5.5GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It feature...$$$openai/gpt-4o-miniGPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more ...$openai/gpt-5.4GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image input...$$openai/gpt-5-miniGPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency an...$openai/gpt-5.4-miniGPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoni...$$openai/gpt-oss-120bgpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B par...$