SandBase is live — $1 in free credits on signupStart free ›

MiniMax modelsaudio generation api

minimax/music/v2

Music V2 by MiniMax - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.

Input
A description of the music, specifying style, mood, and scenario. 10-300 characters.
Lyrics of the song. Use n to separate lines. You may add structure tags like [Intro], [Verse], [Chorus], [Bridge], [Outro] to enhance the arrangement. 10-3000 characters.
Audio configuration settings
Idle

Example output — click Run to generate your own

API README

Minimax Music

Minimax Music turns a written musical brief into a complete generated song or instrumental track. The prompt can define genre, instrumentation, mood, energy, and listening scenario, while a separate lyrics field provides the words and high-level arrangement for vocal music.

The model is designed for composition rather than isolated sound effects or spoken narration. Section labels such as verse, chorus, bridge, and outro help organize longer pieces, making the route useful for demos, concept tracks, campaign music, and early exploration before a production team moves into detailed arrangement.

Highlights

  • Complete text-to-music generation. Creates a structured musical composition from genre, mood, instrumentation, and scenario direction.
  • Lyric-synchronized vocals. Connects supplied words to the vocal performance instead of treating lyrics as unrelated metadata.
  • Multi-section song structure. Uses arrangement labels to shape verses, choruses, bridges, intros, and outros.
  • Broad musical exploration. Supports contrasting genres, emotional arcs, and production directions for rapid concept development.

Pricing

ConfigurationBilling unitPrice
Base generationPer request$0.03

When to Use

✅ Good fit❌ Consider alternatives
The project needs drafting an original songThe goal is a different media task or endpoint
The available inputs match the required local schemaRequired source media or permissions are unavailable
The brief can specify subject, composition, style, and deliveryThe result must be deterministic at pixel or sample level
The supported formats and controls match final placementDelivery requires unsupported dimensions, codecs, or duration
An asynchronous generated result fits the workflowA live, frame-synchronous, or real-time response is mandatory

Prompt Guide

Write the request as a production brief: identify the main subject or source material, state the intended transformation, describe composition or timing, and finish with style, atmosphere, and delivery constraints. Keep preservation requirements separate from requested changes, and use only fields exposed by this route.

{
  "prompt": "A cinematic, precisely directed scene with clear subject action, camera movement, lighting, and atmosphere",
  "audio_setting": "https://example.com/reference.mp3",
  "lyrics_prompt": "example"
}

Technical Specs

SpecValue
Model IDminimax/music/v2
Input fieldsprompt (string)<br>audio_setting (string)<br>lyrics_prompt (string)
Required inputprompt
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

  • minimax/speech/2.8/hd
  • minimax/speech/2.8/turbo
  • minimax/voice-clone

Related Models