Live pricing on OneInfer
Pulled from the OneInfer catalog (OpenRouter and direct provider routes) at request time — rates can change. Capture the timestamp with any benchmark or production decision.
| Model | Provider | Input $/1M tokens | Output $/1M tokens |
|---|---|---|---|
| GLM-5.3-Flash | zai | $140.00 | $440.00 |
| GLM-5.3-Flash | zai | $100.00 | $320.00 |
| GLM-5.3-Flash | zai | $7.50 | $25.00 |
| GLM-5.3-Flash | novita | $7.50 | $25.00 |
| GLM-5.3-Flash | together_ai | $15.00 | $50.00 |
| GLM-5.2 | novita | $140.00 | $440.00 |
| GLM-5.2 | bytedance | $140.00 | $440.00 |
| GLM-5.2 | akashml | $140.00 | $440.00 |
| GLM-5.2 | zai | $140.00 | $440.00 |
Decision snapshot
| Decision factor | GLM-5.3-Flash | GLM-5.2 |
|---|---|---|
| Text and agent workflows | New multimodal flash tier with 26 Aug 2026 release | Prior Z.ai flagship for long-horizon coding and agentic work |
| Input price per 1M tokens | $0.075 | Verify current price on Novita |
| Output price per 1M tokens | $0.250 | Verify current price on Novita |
| Context window | 1,048,576 tokens | 1,000,000 tokens |
| Input modalities | Text, image, video, file | Text only |
| Architecture | Mixture-of-experts (~320B params) | Mixture-of-experts with IndexShare (753B total / 39B active) |
| DeepSWE v1.1 (vendor-reported) | 63.4 | 46.2 |
| Vendor | Z.ai | Z.ai |
| Provider availability | Z.ai first-party | Novita (verify current routing) |
How to read this comparison
GLM-5.3-Flash is not a 5.2 successor at the same tier — it is a new multimodal flash release on top of the Z.ai line. GLM-5.2 is the prior flagship for long-horizon coding and agentic work. The honest framing: this is not "upgrade or downgrade", it is "same vendor, different tier" — and the DeepSWE 63.4 vs 46.2 number is the cleanest single delta Z.ai reports between the two. Treat any "newer is better" framing as marketing and reproduce on your own prompts.
Pricing context
GLM-5.3-Flash lists at $0.075 input / $0.250 output per 1M tokens on the Z.ai / OpenRouter listing as of 27 Aug 2026. GLM-5.2's current price depends on the inference route — the OneInfer catalog lists it via Novita; verify the live rate before assuming parity or premium. At a 1:3 input:output ratio, GLM-5.3-Flash blends to ~$0.21 per 1M tokens; recompute the GLM-5.2 blended rate from its current list price before deciding.
Ready to test the workflow?
Create account & add creditsStrengths and tradeoffs
GLM-5.3-Flash wins on multimodal input (text + image + video + file), newer release date (26 Aug 2026), and a publicly listed flash-tier price. GLM-5.2 wins on long-horizon stability, the IndexShare architecture that cuts per-token FLOPs 2.9× at 1M context, and a track record on existing Z.ai integrations. They are not the same tier — 5.2 is the prior flagship, 5.3-Flash is the new multimodal sibling.
Workload recommendation
| Workload | Recommended starting point |
|---|---|
| Multimodal input (image, video, file) at flash tier | GLM-5.3-Flash |
| Long-horizon agentic coding on existing 5.2 integration | GLM-5.2 (stay until 5.3-Flash validates) |
| Pure text with 1M-token context and FLOP efficiency matters | GLM-5.2 (IndexShare) |
| Vendor-managed multimodal SLA today | GLM-5.3-Flash |
| Reproducible upgrade decision | Run fixed-prompt set on both, compare DeepSWE and your task metrics |
Frequently asked questions
Is GLM-5.3-Flash an upgrade over GLM-5.2?
It is a different tier, not a direct upgrade. GLM-5.3-Flash is a new multimodal flash release (26 Aug 2026) at $0.075 input / $0.250 output per 1M tokens. GLM-5.2 is the prior Z.ai flagship (753B / 39B active, IndexShare architecture) for long-horizon coding. Compare DeepSWE 63.4 vs 46.2 as the cleanest reported delta and validate on your own prompts before switching.
Is GLM-5.2 multimodal like GLM-5.3-Flash?
No — GLM-5.2 accepts text input only. GLM-5.3-Flash adds image, video, and file input at the same 1M context window.
Where can I use GLM-5.2 today?
The OneInfer catalog lists GLM-5.2 via Novita. Verify the current provider routing and live price before relying on it — the 5.2 listing predates the 5.3 launch and provider availability may shift.
Can I try GLM-5.3 before integrating it?
Use the OneInfer GLM-5.3 launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.
How should I treat benchmark claims?
Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.
Put GLM-5.3 to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.