zai

ZAI: GLM 5.3 Flash

z-ai/GLM-5.3-Flash
Create account & add credits
Context

1M

Input

text, image

Output

text

Tool calling

Supported

About this model

GLM-5.3-Flash is Z.AI's natively multimodal Mixture-of-Experts model in the GLM-5 family, designed for high-performance coding, agentic engineering, multimodal reasoning, long-context processing, and efficient inference. It has 320B total parameters with 18B active parameters and introduces a hybrid architecture combining sparse and linear attention together with Manifold-Constrained Hyper-Connections for improved scaling and long-context efficiency. The model supports controllable reasoning effort and is optimized to deliver frontier-level capability at Flash-tier serving cost.

Input modalities

text

Accepted as model input

image

Accepted as model input

Providers

Available routing options for this model through OneInfer.

Available

Input tokens

$0.075 / 1M tokens

Output tokens

$0.250 / 1M tokens

Routing

OneInfer optimized

Available

Input tokens

$0.075 / 1M tokens

Output tokens

$0.250 / 1M tokens

Routing

OneInfer optimized

Available

Input tokens

$0.150 / 1M tokens

Output tokens

$0.500 / 1M tokens

Routing

OneInfer optimized

Pricing

Current OneInfer pricing for this model.

UsagePrice
Input tokens$0.075 / 1M tokens
Output tokens$0.250 / 1M tokens
Cached input tokens$0.015 / 1M tokens

Performance

Published evaluation results associated with this model.

Coding

Terminal-Bench 2.184.3
DeepSWE v1.163.4
NL2Repo56.3

Agentic

Toolathlon Verified78.4
AutomationBench v1.0.648.8
Agents' Last Exam26.3
HLE w/ Tools55.3
GDPval-AA v21773

Vision

OfficeQA Pro62.4
CharXiv Reasoning w/ Tools89.4
Chartography w/ Tools78
BabyVision53.4
MVBench77.8
MMVU80.5

API example

curl https://api.oneinfer.ai/v1/ula/chat/completions \
  -H "Authorization: Bearer $JWT_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/GLM-5.3-Flash",
    "messages": [
      { "role": "user", "content": "Hello!" }
    ]
  }'