Model facet · Benchmarks

Qwen3.8-2.4T-A95B Benchmarks: 86.6 Terminal Bench, 92.6 GPQA

Vendor-reported headline scores from the Hugging Face model card list Terminal Bench 2.1 at 86.6, GPQA Diamond at 92.6, and SWE-bench Pro at 67.7. Five evals are Qwen's own. Independent Artificial Analysis numbers agree closely on GPQA and HLE.

Headline scores — third-party evals

Source: Hugging Face model card, published August 2026. Harness conditions stated per row. Independent scores from Artificial Analysis included where available.

BenchmarkQwen scoreOpus 4.8Fable 5GPT-5.6 SolIndependent (AA)Harness
Terminal Bench 2.186.684.684.688.8Claude Code avg@10, 5-hour timeout, max_tokens=131,072
PaperBench93.090.5Qwen internal, listed as third-party comparison target
GPQA Diamond92.692.092.694.192.7Standard GPQA harness; AA reports the same score within rounding
IFBench82.8Qwen internal
SWE-bench Pro67.769.280.0Terminus 2 / Artificial Analysis for Opus and Fable
DeepSWE 1.156.670.073.0Standard DeepSWE harness
HLE43.653.347.243.0Standard HLE harness; AA score within 0.6
MRCR v2 256K (8-needle)92.9Qwen internal
HealthBench60.255.3Qwen internal
PRBench-Finance58.355.8Qwen internal
PLawBench73.269.672.3Qwen internal
LongBench v266.369.1Standard LongBench v2

In-house Qwen evals — visually separate

Five benchmarks are Qwen's own: QwenSWEBench, QwenQoderBench, QwenReactBench, QwenSVGBench, and CoWorkBench. Treat them as vendor-reported and reproduce them yourself before quoting. CoWorkBench 74.8 is the headline in-house agentic eval; WorkSpaceBench 67.7 and JobBench 53.4 round out the enterprise productivity trio.

Independent measurement: Artificial Analysis

GPQA 92.7%, HLE 43.0%, SciCode 52.9%, LCR 74.3%, Intelligence Index 58.1, Coding Index 71.8. The independent GPQA and HLE scores agree with the vendor-reported numbers within rounding, which is a credible match — most published disagreements on Qwen-Max-class evals sit in the 1–3 point range.

Ready to test the workflow?

Create account & add credits

Measured serving performance

Per cloudprice.net: output at 49.8 tok/s (rank 107), time to first token at 1.80s (rank 489). This is not a low-latency model — the 2.4T scale and required thinking block push output speed into the back half of the cost-vs-latency curve.

Four caveats to state alongside any quoted score

(1) The benchmark table on the official model card is labelled Qwen3.8-Max, not Qwen3.8-2.4T-A95B — Hugging Face structured metadata attaches the same scores to the A95B repo, which makes republishing defensible with a stated note. (2) Harnesses differ per model within the same row — Terminal Bench 2.1 ran Claude Code avg@10 for Qwen, Terminus 2 for Opus and Fable, and Codex for GPT-5.6 Sol. (3) Five benchmarks are in-house. (4) Independent and official numbers disagree slightly; show both.

Frequently asked questions

What is Qwen3.8-2.4T-A95B's Terminal Bench 2.1 score?

86.6 per the Hugging Face model card, evaluated with Claude Code avg@10, a 5-hour timeout, and max_tokens=131,072. Above Claude Opus 4.8 at 84.6 and Claude Fable 5 at 84.6; below GPT-5.6 Sol at 88.8.

Is Qwen3.8-2.4T-A95B good for coding?

Strong on Terminal Bench 2.1 (86.6) and FrontierSWE (73.5). Mid-table on SWE-bench Pro at 67.7 — below Fable 5 at 80.0 and Opus 4.8 at 69.2 — and behind GPT-5.6 Sol and Fable 5 on DeepSWE 1.1 at 56.6.

Where is the live Qwen3.8-2.4T-A95B model page?

The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This page is a focused facet of that entity, not a replacement for it.

How should I treat benchmark or price claims?

Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.

Is Qwen3.8-2.4T-A95B open source?

Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.

Put Qwen3.8-2.4T-A95B to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.