1M
text, image
text
Supported
About this model
GLM-5.3-Flash is Z.AI's natively multimodal Mixture-of-Experts model in the GLM-5 family, designed for high-performance coding, agentic engineering, multimodal reasoning, long-context processing, and efficient inference. It has 320B total parameters with 18B active parameters and introduces a hybrid architecture combining sparse and linear attention together with Manifold-Constrained Hyper-Connections for improved scaling and long-context efficiency. The model supports controllable reasoning effort and is optimized to deliver frontier-level capability at Flash-tier serving cost.
Explore the GLM-5.3 content hub
Input modalities
text
Accepted as model input
image
Accepted as model input
Providers
Available routing options for this model through OneInfer.
Input tokens
$0.075 / 1M tokens
Output tokens
$0.250 / 1M tokens
Routing
OneInfer optimized
Input tokens
$0.075 / 1M tokens
Output tokens
$0.250 / 1M tokens
Routing
OneInfer optimized
Input tokens
$0.150 / 1M tokens
Output tokens
$0.500 / 1M tokens
Routing
OneInfer optimized
Pricing
Current OneInfer pricing for this model.
| Usage | Price |
|---|---|
| Input tokens | $0.075 / 1M tokens |
| Output tokens | $0.250 / 1M tokens |
| Cached input tokens | $0.015 / 1M tokens |
Performance
Published evaluation results associated with this model.
Coding
Agentic
Vision
API Guide
- 1
Generate your access token
Use your API key to generate the JWT access token required by the Models API.
- 2
Add the JWT token
Copy the generated JWT token and replace
JWT_TOKENin the example below. - 3
Check your credits
If your balance is too low, before calling the model.
- 4
Run the API example
Choose your preferred language, copy the example, and send your first model request.
curl https://api.oneinfer.ai/v1/ula/chat/completions \ -H "Authorization: Bearer $JWT_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "model": "z-ai/GLM-5.3-Flash", "messages": [ { "role": "user", "content": "Hello!" } ] }'