zai

ZAI: GLM 5.3 Flash

z-ai/GLM-5.3-Flash
Create account & add credits
Context

1M

Input

text, image

Output

text

Tool calling

Supported

About this model

GLM-5.3-Flash is Z.AI's natively multimodal Mixture-of-Experts model in the GLM-5 family, designed for high-performance coding, agentic engineering, multimodal reasoning, long-context processing, and efficient inference. It has 320B total parameters with 18B active parameters and introduces a hybrid architecture combining sparse and linear attention together with Manifold-Constrained Hyper-Connections for improved scaling and long-context efficiency. The model supports controllable reasoning effort and is optimized to deliver frontier-level capability at Flash-tier serving cost.

Input modalities

text

Accepted as model input

image

Accepted as model input

Providers

Available routing options for this model through OneInfer.

Available

Input tokens

$0.075 / 1M tokens

Output tokens

$0.250 / 1M tokens

Routing

OneInfer optimized

Available

Input tokens

$0.075 / 1M tokens

Output tokens

$0.250 / 1M tokens

Routing

OneInfer optimized

Available

Input tokens

$0.150 / 1M tokens

Output tokens

$0.500 / 1M tokens

Routing

OneInfer optimized

Pricing

Current OneInfer pricing for this model.

UsagePrice
Input tokens$0.075 / 1M tokens
Output tokens$0.250 / 1M tokens
Cached input tokens$0.015 / 1M tokens

Performance

Published evaluation results associated with this model.

Coding

Terminal-Bench 2.184.3
DeepSWE v1.163.4
NL2Repo56.3

Agentic

Toolathlon Verified78.4
AutomationBench v1.0.648.8
Agents' Last Exam26.3
HLE w/ Tools55.3
GDPval-AA v21773

Vision

OfficeQA Pro62.4
CharXiv Reasoning w/ Tools89.4
Chartography w/ Tools78
BabyVision53.4
MVBench77.8
MMVU80.5

API Guide

Refer Documentation
  1. 1

    Generate your access token

    Use your API key to generate the JWT access token required by the Models API.

  2. 2

    Add the JWT token

    Copy the generated JWT token and replace JWT_TOKEN in the example below.

  3. 3

    Check your credits

    If your balance is too low, before calling the model.

  4. 4

    Run the API example

    Choose your preferred language, copy the example, and send your first model request.

    curl https://api.oneinfer.ai/v1/ula/chat/completions \
      -H "Authorization: Bearer $JWT_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "z-ai/GLM-5.3-Flash",
        "messages": [
          { "role": "user", "content": "Hello!" }
        ]
      }'