Skip to content

ERNIE 4.5 21B A3B

POST/v1/chat/completions

A sophisticated text-based Mixture-of-Experts (MoE) model featuring 21B total parameters with 3B activated per token, delivering exceptional multimodal understanding and generation through heterogeneous MoE structures and modality-isolated routing. Supporting an extensive 131K token context length, the model achieves efficient inference via multi-expert parallel collaboration and quantization, while advanced post-training techniques including SFT, DPO, and UPO ensure optimized performance across diverse applications with specialized routing and balancing losses for superior task handling.

Request body

Parameters supported by this model. Values, defaults, and limits are read from the model registry.

stringmodelrequired

Model identifier. Set to baidu/ernie-4.5-21b-a3b.

Default: baidu/ernie-4.5-21b-a3b

array<object>messagesrequired

Conversation messages in system, user, or assistant order.

Optional<integer>max_tokens

Maximum number of tokens the model may generate in the response.

Range: −∞ to 8000

Optional<number>temperature

Sampling temperature. Lower values are more deterministic; higher values are more creative.

Range: 0 to 2

Default: 1

Optional<number>top_p

Nucleus sampling threshold. Use this or temperature, but usually not both.

Range: 0 to 1

Optional<boolean>stream

When true, returns incremental Server-Sent Events instead of one completed response.

Default: false

Optional<array>tools

Tool definitions that the model may call during the response.

Optional<string>tool_choice

Controls whether the model may call a tool and, when supported, which tool it must call.

Optional<array<string>>stop

Sequences that stop generation when the model produces one of them.

Optional<number>presence_penalty

Penalizes tokens that already appeared, encouraging the model to introduce new topics.

Range: -2 to 2

Optional<number>frequency_penalty

Penalizes repeated tokens, reducing repetition in the generated response.

Range: -2 to 2

Optional<integer>seed

Seed used to make sampling more reproducible when the provider supports it.

Optional<integer>top_k

Request parameter supported by this model.

Range: 0 to ∞

Response Schema

Fields returned by this model API response.

array<object>choicesrequired

Generated completion choices.

stringidrequired

Unique chat completion identifier.

stringmodelrequired

Model that generated the response.

Optional<object>usage

Token usage when available.

Model capabilities

array<string>capability_tagsrequired

Capabilities declared by the model registry.

Default: chat, function_calling

integercontext_lengthrequired

Maximum context window accepted by this model.

Default: 120000 tokens

stringexecution_moderequired

Execution mode declared by the model registry.

Default: sync