GLM 5.3 Flash
z-ai/glm-5.3-flashGLM-5.3-Flash is the first native multimodal model in Z.ai's GLM-5 series, a 320B-parameter mixture-of-experts model with 18B active parameters that combines sparse and linear attention to cut long-context compute and KV cache. Vision is built into its coding loop, letting it inspect rendered interfaces and GUI feedback for frontend, game and computer-use work, and it also handles office documents and financial research. Z.ai reports it outperforms GLM-5.2; GLM-5.3-FlashX serves it at higher speed.
- Input price
- $0.11
$0.15USD / 1M tokens - Output price
- $0.38
$0.50USD / 1M tokens - Context window
- 1M
- Max output
- 131.1K
Try the model
Playground
Try a prompt
Your response will appear here
Choose an example or write a prompt, then click Run.
Model card
Specifications
Pricing
- Input
- $0.11 / 1M tokens
- Output
- $0.38 / 1M tokens
- Cache read
- $0.02 / 1M tokens
Context & modalities
- Context window
- 1,048,576 tokens
- Max output
- 131,072 tokens
- Input
- Text, image
- Output
- Text
Capabilities
- Chat
- Supported
- Vision
- Supported
- Reasoning
- Supported
- Structured output
- Supported
- Function calling
- Supported
- Audio input
- Not supported
Access
- Provider
- Z.ai
- Model ID
- z-ai/glm-5.3-flash
- Execution
- sync
- Base URL
- https://api.sandbase.ai
- API
- Chat Completions API
- Endpoint
- /v1/chat/completions
Independent evaluations
Benchmarks
Scores on standardized evaluations, compared with other models on SandBase.
Metrics sourced from Artificial Analysis · Intelligence Index v4.3
Start building
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
# The quoted heredoc keeps Unicode and shell metacharacters unchanged.
result=$(curl --fail-with-body --silent \
-X POST "https://api.sandbase.ai/v1/chat/completions" \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @- <<'SANDBASE_JSON'
{
"model": "z-ai/glm-5.3-flash",
"messages": [
{
"role": "user",
"content": "Hello"
}
],
"max_tokens": 2048
}
SANDBASE_JSON
)
printf '%s\n' "$result"Choose your model
Compare models
| Model | Context | Input / 1M | Output / 1M | Released |
|---|---|---|---|---|
GLM 5.3 FlashThis model Z.ai | 1M | $0.11 | $0.38 | Aug 26, 2026 |
Z.ai | — | Coming soon | Coming soon | Coming soon |
Z.ai | 1M | $0.37 | $1.25 | Sep 18, 2026 |
Z.ai | 1M | $1.40 | $4.40 | Aug 18, 2026 |
Z.ai | 1M | $1.40 | $4.40 | Jun 16, 2026 |
Z.ai | 128K | $0.20 | $1.10 | Jul 25, 2025 |
Questions
FAQ
How do I call Z.ai GLM 5.3 Flash through SandBase?
Create a SandBase API key, then send requests with the model ID "z-ai/glm-5.3-flash" to Chat Completions API (/v1/chat/completions) at https://api.sandbase.ai. The request examples on this page show the exact payload.
How much does Z.ai GLM 5.3 Flash cost?
$0.11 per 1M input tokens and $0.38 per 1M output tokens with the current 25% discount (list price $0.15 / $0.50), billed per request from your SandBase balance.
What is the context window of Z.ai GLM 5.3 Flash?
1M tokens of context, with up to 131.1K output tokens per response.
How does Z.ai GLM 5.3 Flash score on benchmarks?
Z.ai GLM 5.3 Flash scores 41.8 on the Artificial Analysis Intelligence Index (v4.3), ranking #18 of 84 models on SandBase; Coding Index 71.5, Agentic Index 50.9. Source: Artificial Analysis, measured as GLM 5.3 Flash.
Which features does Z.ai GLM 5.3 Flash support?
Z.ai GLM 5.3 Flash supports vision (image input), reasoning, function calling, structured output. See Specifications above for the full list.
Do I need a separate Z.ai account?
No. One SandBase API key and balance gives you access to Z.ai GLM 5.3 Flash and the other models in the catalog; you do not need to sign up with Z.ai separately.
