Vendor-reported benchmarks
Scores below are sourced from the OneInfer model catalog entry for `z-ai/GLM-5.3-Flash`, 29 Aug 2026. They are vendor-reported rather than independently reproduced — verify against an independent harness (Artificial Analysis, Hugging Face OpenLLM) before relying on any single number.
| Evaluation | Score |
|---|---|
| Coding | |
| Terminal-Bench 2.1 | 84.3 |
| DeepSWE v1.1 | 63.4 |
| NL2Repo | 56.3 |
| Agentic | |
| Toolathlon Verified | 78.4 |
| AutomationBench v1.0.6 | 48.8 |
| Agents' Last Exam | 26.3 |
| HLE w/ Tools | 55.3 |
| GDPval-AA v2 | 1773 |
| Vision | |
| OfficeQA Pro | 62.4 |
| CharXiv Reasoning w/ Tools | 89.4 |
| Chartography w/ Tools | 78 |
| BabyVision | 53.4 |
| MVBench | 77.8 |
| MMVU | 80.5 |
How these scores are sourced
Scores come from `benchmark_info` on the `z-ai/GLM-5.3-Flash` catalog entry. Harness version, sample count, and reproduction conditions are not published in the catalog entry — treat them as vendor-reported until an independent reproduction is published.
Ready to test the workflow?
Create account & add creditsSee the full matrix
When independent reproductions are published, the canonical benchmark matrix will move to /compare/glm-5-3-flash-benchmarks — not duplicated here.
Frequently asked questions
Where is the live GLM-5.3-Flash model page?
The canonical model page with current OneInfer pricing, capabilities, and availability is /models/zai-org/GLM-5.3-Flash. This page is a focused facet of that entity, not a replacement for it.
How should I treat benchmark or price claims?
Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.
Put GLM-5.3-Flash to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.