Model comparison

Qwen3.8-Flash vs DeepSeek V4

Compare coding reliability, serving availability, and data-control requirements. This page separates independently verified results from vendor-reported claims and avoids declaring a universal winner.

Provider pricing (live)

Pulled from the OneInfer pricing catalog; rates can change. Capture the timestamp with any benchmark or production decision.

ProviderInput $ / 1M tokensOutput $ / 1M tokens
openrouter$0.150$0.470

Vendor-reported benchmarks

Scores come from the Qwen/Qwen3.8-27B benchmark record in the OneInfer model catalog and have not been independently reproduced.

EvaluationScore
General agentic
Claw-Eval Avg72.4
Claw-Eval Pass^360.6
QwenClawBench53.4
SkillsBench48.2
Agentic coding
SWE-Bench Verified77.2
SWE-Bench Pro53.5
SWE-Bench Multilingual71.3
TerminalBench 2.059.3
NL2Repo36.2
QwenWebBench1487.0
Multimodal
MMMU82.9
MMMU-Pro75.8
MathVista mini87.4
DynaMath85.6
VlmsAreBlind97.0
RealWorldQA84.1
MMStar81.4
MMBench EN-DEV v1.192.3
SimpleVQA56.1
General capabilities and reasoning
MMLU-Pro86.2
MMLU-Redux93.5
SuperGPQA66.0
C-Eval91.4
GPQA Diamond87.8
Humanity's Last Exam24.0
LiveCodeBench v683.9
AIME 202694.1
HMMT Feb 202684.3
HMMT Nov 202590.7
IMOAnswerBench80.8
Document understanding
CharXiv RQ78.4
CC-OCR81.2
OCRBench89.4
Spatial intelligence
ERQA62.5
CountBench97.8
RefCOCO Avg92.5
EmbSpatialBench84.6
RefSpatialBench70.0
Video understanding
VideoMME87.7
VideoMMMU84.4
MLVU86.6
MVBench75.5
Visual agent
V*94.7
AndroidWorld70.3

Decision snapshot

Decision factorQwen3.8-FlashDeepSeek V4
Text and agent workflowsEvaluateEvaluate
BenchmarksVendor-reported; live aboveVerify independent evaluation
Deployment controlVerify weight/license statusVerify current terms
Non-text modalitiesVerify provider capabilitiesVerify provider capabilities
Input $ / 1M tokens$0.16 (live above)$0.66 ($0.14 V4 Flash)
Output $ / 1M tokens$0.47 (live above)$1.98 ($0.28 V4 Flash)
Headline benchmarkVendor-reported; live aboveDeepSeek V4 vendor release notes: DeepSeek vendor-reported (api-docs.deepseek.com/news/news260424)
Pricing provenanceDeepSeek platform — https://platform.deepseek.com/

Rival pricing anchor (vendor-reported)

DeepSeek V4 (1.6T MoE, ~49B active) at off-peak rates; peak tier is cheaper (~$0.079 / $0.278). V4 Flash is a 284B MoE (~13B active). 1M context.

Ready to test the workflow?

Create account & add credits

Benchmark rules

  • Compare only the same evaluation and harness version.
  • Label vendor-reported results.
  • Record token budget and tool policy.
  • Do not infer production reliability from one benchmark.

Strengths and tradeoffs

Select the model against a representative prompt set, latency target, output budget, tool-calling requirements, and data-control constraints.

Workload recommendation

WorkloadHow to choose
Budget-sensitive codingCompare task success per dollar.
Long-context analysisTest retrieval and citation accuracy.
Multimodal inputChoose a model/provider that explicitly supports it.
Regulated dataReview retention, residency, and deployment terms.

Frequently asked questions

Can I try Qwen3.8-Flash before integrating it?

Use the OneInfer Qwen3.8-Flash launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.

How should I treat benchmark claims?

Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.

Put Qwen3.8-Flash to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.