Model comparison

Qwen3.8-Flash benchmark matrix

Use this matrix as an evidence index, not a single composite verdict. This page separates independently verified results from vendor-reported claims and avoids declaring a universal winner.

Provider pricing (live)

Pulled from the OneInfer pricing catalog; rates can change. Capture the timestamp with any benchmark or production decision.

ProviderInput $ / 1M tokensOutput $ / 1M tokens
openrouter$0.150$0.470

Vendor-reported benchmarks

Scores come from the Qwen/Qwen3.8-27B benchmark record in the OneInfer model catalog and have not been independently reproduced.

EvaluationScore
General agentic
Claw-Eval Avg72.4
Claw-Eval Pass^360.6
QwenClawBench53.4
SkillsBench48.2
Agentic coding
SWE-Bench Verified77.2
SWE-Bench Pro53.5
SWE-Bench Multilingual71.3
TerminalBench 2.059.3
NL2Repo36.2
QwenWebBench1487.0
Multimodal
MMMU82.9
MMMU-Pro75.8
MathVista mini87.4
DynaMath85.6
VlmsAreBlind97.0
RealWorldQA84.1
MMStar81.4
MMBench EN-DEV v1.192.3
SimpleVQA56.1
General capabilities and reasoning
MMLU-Pro86.2
MMLU-Redux93.5
SuperGPQA66.0
C-Eval91.4
GPQA Diamond87.8
Humanity's Last Exam24.0
LiveCodeBench v683.9
AIME 202694.1
HMMT Feb 202684.3
HMMT Nov 202590.7
IMOAnswerBench80.8
Document understanding
CharXiv RQ78.4
CC-OCR81.2
OCRBench89.4
Spatial intelligence
ERQA62.5
CountBench97.8
RefCOCO Avg92.5
EmbSpatialBench84.6
RefSpatialBench70.0
Video understanding
VideoMME87.7
VideoMMMU84.4
MLVU86.6
MVBench75.5
Visual agent
V*94.7
AndroidWorld70.3

Provenance-first matrix

Live benchmark records appear here only with evaluation name, harness version, source, and verification date.

Ready to test the workflow?

Create account & add credits

Compatibility rule

Scores from different harness versions, tool policies, or token budgets are not directly comparable.

Frequently asked questions

Can I try Qwen3.8-Flash before integrating it?

Use the OneInfer Qwen3.8-Flash launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.

How should I treat benchmark claims?

Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.

Put Qwen3.8-Flash to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.