1M
text, image
text
Supported
About this model
GLM-5.3-Flash is Z.AI's natively multimodal Mixture-of-Experts model in the GLM-5 family, designed for high-performance coding, agentic engineering, multimodal reasoning, long-context processing, and efficient inference. It has 320B total parameters with 18B active parameters and introduces a hybrid architecture combining sparse and linear attention together with Manifold-Constrained Hyper-Connections for improved scaling and long-context efficiency. The model supports controllable reasoning effort and is optimized to deliver frontier-level capability at Flash-tier serving cost.
Explore the GLM-5.3 content hub
Input modalities
text
Accepted as model input
image
Accepted as model input
Providers
Available routing options for this model through OneInfer.
Input tokens
$0.075 / 1M tokens
Output tokens
$0.250 / 1M tokens
Routing
OneInfer optimized
Input tokens
$0.075 / 1M tokens
Output tokens
$0.250 / 1M tokens
Routing
OneInfer optimized
Input tokens
$0.150 / 1M tokens
Output tokens
$0.500 / 1M tokens
Routing
OneInfer optimized
Pricing
Current OneInfer pricing for this model.
| Usage | Price |
|---|---|
| Input tokens | $0.075 / 1M tokens |
| Output tokens | $0.250 / 1M tokens |
| Cached input tokens | $0.015 / 1M tokens |
Performance
Published evaluation results associated with this model.
Coding
Agentic
Vision
API example
curl https://api.oneinfer.ai/v1/ula/chat/completions \
-H "Authorization: Bearer $JWT_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "z-ai/GLM-5.3-Flash",
"messages": [
{ "role": "user", "content": "Hello!" }
]
}'