Provider pricing (live)
Pulled from the OneInfer pricing catalog; rates can change. Capture the timestamp with any benchmark or production decision.
| Provider | Input $ / 1M tokens | Output $ / 1M tokens |
|---|---|---|
| openrouter | $0.150 | $0.470 |
Vendor-reported benchmarks
Scores come from the Qwen/Qwen3.8-27B benchmark record in the OneInfer model catalog and have not been independently reproduced.
| Evaluation | Score |
|---|---|
| General agentic | |
| Claw-Eval Avg | 72.4 |
| Claw-Eval Pass^3 | 60.6 |
| QwenClawBench | 53.4 |
| SkillsBench | 48.2 |
| Agentic coding | |
| SWE-Bench Verified | 77.2 |
| SWE-Bench Pro | 53.5 |
| SWE-Bench Multilingual | 71.3 |
| TerminalBench 2.0 | 59.3 |
| NL2Repo | 36.2 |
| QwenWebBench | 1487.0 |
| Multimodal | |
| MMMU | 82.9 |
| MMMU-Pro | 75.8 |
| MathVista mini | 87.4 |
| DynaMath | 85.6 |
| VlmsAreBlind | 97.0 |
| RealWorldQA | 84.1 |
| MMStar | 81.4 |
| MMBench EN-DEV v1.1 | 92.3 |
| SimpleVQA | 56.1 |
| General capabilities and reasoning | |
| MMLU-Pro | 86.2 |
| MMLU-Redux | 93.5 |
| SuperGPQA | 66.0 |
| C-Eval | 91.4 |
| GPQA Diamond | 87.8 |
| Humanity's Last Exam | 24.0 |
| LiveCodeBench v6 | 83.9 |
| AIME 2026 | 94.1 |
| HMMT Feb 2026 | 84.3 |
| HMMT Nov 2025 | 90.7 |
| IMOAnswerBench | 80.8 |
| Document understanding | |
| CharXiv RQ | 78.4 |
| CC-OCR | 81.2 |
| OCRBench | 89.4 |
| Spatial intelligence | |
| ERQA | 62.5 |
| CountBench | 97.8 |
| RefCOCO Avg | 92.5 |
| EmbSpatialBench | 84.6 |
| RefSpatialBench | 70.0 |
| Video understanding | |
| VideoMME | 87.7 |
| VideoMMMU | 84.4 |
| MLVU | 86.6 |
| MVBench | 75.5 |
| Visual agent | |
| V* | 94.7 |
| AndroidWorld | 70.3 |
Decision snapshot
| Decision factor | Qwen3.8-Flash | GLM-5.3 |
|---|---|---|
| Text and agent workflows | Evaluate | Evaluate |
| Benchmarks | Vendor-reported; live above | Verify independent evaluation |
| Deployment control | Verify weight/license status | Verify current terms |
| Non-text modalities | Verify provider capabilities | Verify provider capabilities |
| Input $ / 1M tokens | $0.16 (live above) | $1.40 ($0.15 Flash) |
| Output $ / 1M tokens | $0.47 (live above) | $4.40 ($0.50 Flash) |
| Headline benchmark | Vendor-reported; live above | GLM-5.3 frontier coding + emergent cyber capabilities: Z.ai vendor-reported (z.ai/blog/glm-5.3) |
| Pricing provenance | Z.ai GLM pricing — https://docs.z.ai/guides/pricing/pricing |
Rival pricing anchor (vendor-reported)
Z.ai ships GLM-5.3 (full, $1.40 / $4.40) and GLM-5.3-Flash ($0.15 / $0.50). GLM-5.3-Flash is a 320B-A18B MoE with 1M context.
Ready to test the workflow?
Create account & add creditsBenchmark rules
- Compare only the same evaluation and harness version.
- Label vendor-reported results.
- Record token budget and tool policy.
- Do not infer production reliability from one benchmark.
Strengths and tradeoffs
Select the model against a representative prompt set, latency target, output budget, tool-calling requirements, and data-control constraints.
Workload recommendation
| Workload | How to choose |
|---|---|
| Budget-sensitive coding | Compare task success per dollar. |
| Long-context analysis | Test retrieval and citation accuracy. |
| Multimodal input | Choose a model/provider that explicitly supports it. |
| Regulated data | Review retention, residency, and deployment terms. |
Frequently asked questions
Can I try Qwen3.8-Flash before integrating it?
Use the OneInfer Qwen3.8-Flash launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.
How should I treat benchmark claims?
Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.
Put Qwen3.8-Flash to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.