Headline scores — third-party evals
Source: Hugging Face model card, published August 2026. Harness conditions stated per row. Independent scores from Artificial Analysis included where available.
| Benchmark | Qwen score | Opus 4.8 | Fable 5 | GPT-5.6 Sol | Independent (AA) | Harness |
|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 86.6 | 84.6 | 84.6 | 88.8 | — | Claude Code avg@10, 5-hour timeout, max_tokens=131,072 |
| PaperBench | 93.0 | — | — | 90.5 | — | Qwen internal, listed as third-party comparison target |
| GPQA Diamond | 92.6 | 92.0 | 92.6 | 94.1 | 92.7 | Standard GPQA harness; AA reports the same score within rounding |
| IFBench | 82.8 | — | — | — | — | Qwen internal |
| SWE-bench Pro | 67.7 | 69.2 | 80.0 | — | — | Terminus 2 / Artificial Analysis for Opus and Fable |
| DeepSWE 1.1 | 56.6 | — | 70.0 | 73.0 | — | Standard DeepSWE harness |
| HLE | 43.6 | — | 53.3 | 47.2 | 43.0 | Standard HLE harness; AA score within 0.6 |
| MRCR v2 256K (8-needle) | 92.9 | — | — | — | — | Qwen internal |
| HealthBench | 60.2 | — | — | 55.3 | — | Qwen internal |
| PRBench-Finance | 58.3 | — | 55.8 | — | — | Qwen internal |
| PLawBench | 73.2 | 69.6 | — | 72.3 | — | Qwen internal |
| LongBench v2 | 66.3 | 69.1 | — | — | — | Standard LongBench v2 |
In-house Qwen evals — visually separate
Five benchmarks are Qwen's own: QwenSWEBench, QwenQoderBench, QwenReactBench, QwenSVGBench, and CoWorkBench. Treat them as vendor-reported and reproduce them yourself before quoting. CoWorkBench 74.8 is the headline in-house agentic eval; WorkSpaceBench 67.7 and JobBench 53.4 round out the enterprise productivity trio.
Independent measurement: Artificial Analysis
GPQA 92.7%, HLE 43.0%, SciCode 52.9%, LCR 74.3%, Intelligence Index 58.1, Coding Index 71.8. The independent GPQA and HLE scores agree with the vendor-reported numbers within rounding, which is a credible match — most published disagreements on Qwen-Max-class evals sit in the 1–3 point range.
Ready to test the workflow?
Create account & add creditsMeasured serving performance
Per cloudprice.net: output at 49.8 tok/s (rank 107), time to first token at 1.80s (rank 489). This is not a low-latency model — the 2.4T scale and required thinking block push output speed into the back half of the cost-vs-latency curve.
Four caveats to state alongside any quoted score
(1) The benchmark table on the official model card is labelled Qwen3.8-Max, not Qwen3.8-2.4T-A95B — Hugging Face structured metadata attaches the same scores to the A95B repo, which makes republishing defensible with a stated note. (2) Harnesses differ per model within the same row — Terminal Bench 2.1 ran Claude Code avg@10 for Qwen, Terminus 2 for Opus and Fable, and Codex for GPT-5.6 Sol. (3) Five benchmarks are in-house. (4) Independent and official numbers disagree slightly; show both.
Frequently asked questions
What is Qwen3.8-2.4T-A95B's Terminal Bench 2.1 score?
86.6 per the Hugging Face model card, evaluated with Claude Code avg@10, a 5-hour timeout, and max_tokens=131,072. Above Claude Opus 4.8 at 84.6 and Claude Fable 5 at 84.6; below GPT-5.6 Sol at 88.8.
Is Qwen3.8-2.4T-A95B good for coding?
Strong on Terminal Bench 2.1 (86.6) and FrontierSWE (73.5). Mid-table on SWE-bench Pro at 67.7 — below Fable 5 at 80.0 and Opus 4.8 at 69.2 — and behind GPT-5.6 Sol and Fable 5 on DeepSWE 1.1 at 56.6.
Where is the live Qwen3.8-2.4T-A95B model page?
The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This page is a focused facet of that entity, not a replacement for it.
How should I treat benchmark or price claims?
Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.
Is Qwen3.8-2.4T-A95B open source?
Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.
Put Qwen3.8-2.4T-A95B to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.