Tracked providers
| Provider | Input /1M tokens | Output /1M tokens | Provenance |
|---|---|---|---|
| Z.ai | $0.075 | $0.250 | vendor reported |
| Novita | $0.075 | $0.250 | vendor reported |
| Together AI | $0.150 | $0.500 | vendor reported |
What is GLM-5.3-Flash?
GLM-5.3-Flash is Z.ai's natively multimodal mixture-of-experts model in the GLM-5 family — 320B total parameters with 18B active per token, BF16 weights, 1,048,576-token context window, image + text input, text output. Released 26 Aug 2026, priced at $0.075 per 1M input tokens and $0.250 per 1M output tokens on the OpenRouter listing.
Providers
Each row is labeled by provenance: independently verified listings versus a provider's own vendor-reported number. Routing, quantization, and uptime can all change — record the provider and date alongside any benchmark or latency claim you rely on.
Ready to test the workflow?
Create account & add creditsWhen GLM-5.3-Flash is not the right choice
GLM-5.3-Flash is not the cheapest flash-tier option on input or output — DeepSeek V4 Flash 0731 is 2.5× cheaper on input and 3.3× cheaper on output. The defensive position for GLM-5.3-Flash is native multimodal (image + text input) at flash-tier pricing. If you need only text, DeepSeek V4 Flash 0731 is the cheaper option; if you need stronger reasoning on coding agents, the Qwen3.8-Max tier is the upgrade path.
Frequently asked questions
Where is the live GLM-5.3-Flash model page?
The canonical model page with current OneInfer pricing, capabilities, and availability is /models/zai-org/GLM-5.3-Flash. This page is a focused facet of that entity, not a replacement for it.
How should I treat benchmark or price claims?
Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.
Put GLM-5.3-Flash to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.