Vendor-reported benchmarks
Scores as published by the vendor on the OneInfer model catalog — not independently reproduced.
| Evaluation | Score |
|---|---|
| General agentic | |
| Claw-Eval Avg | 72.4 |
| Claw-Eval Pass^3 | 60.6 |
| QwenClawBench | 53.4 |
| SkillsBench | 48.2 |
| Agentic coding | |
| SWE-Bench Verified | 77.2 |
| SWE-Bench Pro | 53.5 |
| SWE-Bench Multilingual | 71.3 |
| TerminalBench 2.0 | 59.3 |
| NL2Repo | 36.2 |
| QwenWebBench | 1487.0 |
| Multimodal | |
| MMMU | 82.9 |
| MMMU-Pro | 75.8 |
| MathVista mini | 87.4 |
| DynaMath | 85.6 |
| VlmsAreBlind | 97.0 |
| RealWorldQA | 84.1 |
| MMStar | 81.4 |
| MMBench EN-DEV v1.1 | 92.3 |
| SimpleVQA | 56.1 |
| General capabilities and reasoning | |
| MMLU-Pro | 86.2 |
| MMLU-Redux | 93.5 |
| SuperGPQA | 66.0 |
| C-Eval | 91.4 |
| GPQA Diamond | 87.8 |
| Humanity's Last Exam | 24.0 |
| LiveCodeBench v6 | 83.9 |
| AIME 2026 | 94.1 |
| HMMT Feb 2026 | 84.3 |
| HMMT Nov 2025 | 90.7 |
| IMOAnswerBench | 80.8 |
| Document understanding | |
| CharXiv RQ | 78.4 |
| CC-OCR | 81.2 |
| OCRBench | 89.4 |
| Spatial intelligence | |
| ERQA | 62.5 |
| CountBench | 97.8 |
| RefCOCO Avg | 92.5 |
| EmbSpatialBench | 84.6 |
| RefSpatialBench | 70.0 |
| Video understanding | |
| VideoMME | 87.7 |
| VideoMMMU | 84.4 |
| MLVU | 86.6 |
| MVBench | 75.5 |
| Visual agent | |
| V* | 94.7 |
| AndroidWorld | 70.3 |
See the full matrix
When evaluations are populated, the canonical benchmark matrix will be at /compare/qwen3-8-flash-benchmarks — not duplicated here.
Ready to test the workflow?
Create account & add creditsFrequently asked questions
Where is the live Qwen3.8-Flash model page?
The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-Flash. This page is a focused facet of that entity, not a replacement for it.
How should I treat benchmark or price claims?
Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.
Put Qwen3.8-Flash to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.