Skip to content

ACE-Step

POST/v1/run

Ace Step by ace - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.

Request body

Submit an async generation request. The model field selects the model; other fields are model-specific input parameters.

stringmodelrequired

Model identifier. Set to ace/ace-step.

Default: ace/ace-step

Optional<string>lyrics

Lyrics to be sung in the audio. If not provided or if [inst] or [instrumental] is the content of this field, no lyrics will be sung. Use control structures like [verse], [chorus] and [bridge] to control the structure of the song.

Default:

Optional<integer>duration

The duration of the generated audio in seconds.

Range: 5 to 240

Default: 60

Optional<number>guidance_scale

Guidance scale for the generation.

Range: 0 to 200

Default: 15

Optional<integer>seed

Random seed for reproducibility. If not provided, a random seed will be used.

Optional<number>lyric_guidance_scale

Lyric guidance scale for the generation.

Range: 0 to 10

Default: 1.5

Optional<number>tag_guidance_scale

Tag guidance scale for the generation.

Range: 0 to 10

Default: 5

Optional<string>guidance_type

Type of CFG to use for the generation process.

Allowed values: cfg, apg, cfg_star

Default: apg

Optional<string>scheduler

Scheduler to use for the generation process.

Allowed values: euler, heun

Default: euler

Optional<number>guidance_interval

Guidance interval for the generation. 0.5 means only apply guidance in the middle steps (0.25 * infer_steps to 0.75 * infer_steps)

Range: 0 to 1

Default: 0.5

Optional<number>guidance_interval_decay

Guidance interval decay for the generation. Guidance scale will decay from guidance_scale to min_guidance_scale in the interval. 0.0 means no decay.

Range: 0 to 1

Default: 0

Optional<number>minimum_guidance_scale

Minimum guidance scale for the generation after the decay.

Range: 0 to 200

Default: 3

Optional<integer>granularity_scale

Granularity scale for the generation process. Higher values can reduce artifacts.

Range: -100 to 100

Default: 10

Optional<integer>number_of_steps

Number of steps to generate the audio.

Range: 3 to 60

Default: 27

Optional<string>tags

Comma-separated list of genre tags to control the style of the generated audio.

Response Schema

The submit endpoint returns a run response. If its status is pending or running, poll GET /v1/run/{id} with the returned opaque ID until it reaches a terminal state.

stringidrequired

Opaque SandBase run identifier. Use it exactly as returned; no prefix is guaranteed.

stringstatusrequired

Current public run status.

Allowed values: pending, running, completed, failed, timeout

Optional<string>model

Public SandBase model name used for this run.

Optional<array<object>>outputs

Present only for completed runs. Each object is capability-specific; inspect the selected model schema for its fields.

Optional<object>error

Present only for failed or timeout runs. Contains a public error type and sanitized message.

Optional<string>error.type

Stable public error category.

Optional<string>error.message

Sanitized error message safe to show to clients.

Optional<object>usage

Usage details when available.

Model capabilities

array<string>capability_tagsrequired

Capabilities declared by the model registry.

Default: text-to-audio

stringexecution_moderequired

Execution mode declared by the model registry.

Default: async