Skip to content

ACE-Step

POST/v1/run

Ace Step Audio To Audio by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Request body

Submit an async generation request. The model field selects the model; other fields are model-specific input parameters.

stringmodelrequired

Model identifier. Set to ace/ace-step/audio-to-audio.

Default: ace/ace-step/audio-to-audio

Optional<string>audio

URL of the audio file to be outpainted.

Optional<string>lyrics

Lyrics to be sung in the audio. If not provided or if [inst] or [instrumental] is the content of this field, no lyrics will be sung. Use control structures like [verse], [chorus] and [bridge] to control the structure of the song.

Default:

Optional<number>guidance_scale

Guidance scale for the generation.

Range: 0 to 200

Default: 15

Optional<integer>seed

Random seed for reproducibility. If not provided, a random seed will be used.

Optional<number>minimum_guidance_scale

Minimum guidance scale for the generation after the decay.

Range: 0 to 200

Default: 3

Optional<number>tag_guidance_scale

Tag guidance scale for the generation.

Range: 0 to 10

Default: 5

Optional<number>lyric_guidance_scale

Lyric guidance scale for the generation.

Range: 0 to 10

Default: 1.5

Optional<string>edit_mode

Whether to edit the lyrics only or remix the audio.

Allowed values: lyrics, remix

Default: remix

Optional<number>guidance_interval

Guidance interval for the generation. 0.5 means only apply guidance in the middle steps (0.25 * infer_steps to 0.75 * infer_steps)

Range: 0 to 1

Default: 0.5

Optional<integer>original_seed

Original seed of the audio file.

Optional<string>scheduler

Scheduler to use for the generation process.

Allowed values: euler, heun

Default: euler

Optional<integer>granularity_scale

Granularity scale for the generation process. Higher values can reduce artifacts.

Range: -100 to 100

Default: 10

Optional<string>guidance_type

Type of CFG to use for the generation process.

Allowed values: cfg, apg, cfg_star

Default: apg

Optional<string>original_lyrics

Original lyrics of the audio file.

Default:

Optional<string>original_tags

Original tags of the audio file.

Optional<number>guidance_interval_decay

Guidance interval decay for the generation. Guidance scale will decay from guidance_scale to min_guidance_scale in the interval. 0.0 means no decay.

Range: 0 to 1

Default: 0

Optional<integer>number_of_steps

Number of steps to generate the audio.

Range: 3 to 60

Default: 27

Optional<string>tags

Comma-separated list of genre tags to control the style of the generated audio.

Response Schema

The submit endpoint returns a run response. If its status is pending or running, poll GET /v1/run/{id} with the returned opaque ID until it reaches a terminal state.

stringidrequired

Opaque SandBase run identifier. Use it exactly as returned; no prefix is guaranteed.

stringstatusrequired

Current public run status.

Allowed values: pending, running, completed, failed, timeout

Optional<string>model

Public SandBase model name used for this run.

Optional<array<object>>outputs

Present only for completed runs. Each object is capability-specific; inspect the selected model schema for its fields.

Optional<object>error

Present only for failed or timeout runs. Contains a public error type and sanitized message.

Optional<string>error.type

Stable public error category.

Optional<string>error.message

Sanitized error message safe to show to clients.

Optional<object>usage

Usage details when available.

Model capabilities

array<string>capability_tagsrequired

Capabilities declared by the model registry.

Default: audio-to-audio

stringexecution_moderequired

Execution mode declared by the model registry.

Default: async