Model facet · Benchmarks

GLM-5.3-Flash benchmarks

GLM-5.3-Flash scores are published here from the OneInfer model catalog entry (vendor-reported, 29 Aug 2026) — Coding, Agentic, and Vision evaluations are tracked separately. The full provenance-tagged matrix lives on this page so the exact-match query has a direct answer.

Vendor-reported benchmarks

Scores below are sourced from the OneInfer model catalog entry for `z-ai/GLM-5.3-Flash`, 29 Aug 2026. They are vendor-reported rather than independently reproduced — verify against an independent harness (Artificial Analysis, Hugging Face OpenLLM) before relying on any single number.

EvaluationScore
Coding
Terminal-Bench 2.184.3
DeepSWE v1.163.4
NL2Repo56.3
Agentic
Toolathlon Verified78.4
AutomationBench v1.0.648.8
Agents' Last Exam26.3
HLE w/ Tools55.3
GDPval-AA v21773
Vision
OfficeQA Pro62.4
CharXiv Reasoning w/ Tools89.4
Chartography w/ Tools78
BabyVision53.4
MVBench77.8
MMVU80.5

How these scores are sourced

Scores come from `benchmark_info` on the `z-ai/GLM-5.3-Flash` catalog entry. Harness version, sample count, and reproduction conditions are not published in the catalog entry — treat them as vendor-reported until an independent reproduction is published.

Ready to test the workflow?

Create account & add credits

See the full matrix

When independent reproductions are published, the canonical benchmark matrix will move to /compare/glm-5-3-flash-benchmarks — not duplicated here.

Frequently asked questions

Where is the live GLM-5.3-Flash model page?

The canonical model page with current OneInfer pricing, capabilities, and availability is /models/zai-org/GLM-5.3-Flash. This page is a focused facet of that entity, not a replacement for it.

How should I treat benchmark or price claims?

Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.

Put GLM-5.3-Flash to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.