Provider pricing (live)
Pulled from the OneInfer pricing catalog; rates can change. Capture the timestamp with any benchmark or production decision.
| Provider | Input $ / 1M tokens | Output $ / 1M tokens |
|---|---|---|
| openrouter | $0.150 | $0.470 |
Vendor-reported benchmarks
Scores come from the Qwen/Qwen3.8-27B benchmark record in the OneInfer model catalog and have not been independently reproduced.
| Evaluation | Score |
|---|---|
| General agentic | |
| Claw-Eval Avg | 72.4 |
| Claw-Eval Pass^3 | 60.6 |
| QwenClawBench | 53.4 |
| SkillsBench | 48.2 |
| Agentic coding | |
| SWE-Bench Verified | 77.2 |
| SWE-Bench Pro | 53.5 |
| SWE-Bench Multilingual | 71.3 |
| TerminalBench 2.0 | 59.3 |
| NL2Repo | 36.2 |
| QwenWebBench | 1487.0 |
| Multimodal | |
| MMMU | 82.9 |
| MMMU-Pro | 75.8 |
| MathVista mini | 87.4 |
| DynaMath | 85.6 |
| VlmsAreBlind | 97.0 |
| RealWorldQA | 84.1 |
| MMStar | 81.4 |
| MMBench EN-DEV v1.1 | 92.3 |
| SimpleVQA | 56.1 |
| General capabilities and reasoning | |
| MMLU-Pro | 86.2 |
| MMLU-Redux | 93.5 |
| SuperGPQA | 66.0 |
| C-Eval | 91.4 |
| GPQA Diamond | 87.8 |
| Humanity's Last Exam | 24.0 |
| LiveCodeBench v6 | 83.9 |
| AIME 2026 | 94.1 |
| HMMT Feb 2026 | 84.3 |
| HMMT Nov 2025 | 90.7 |
| IMOAnswerBench | 80.8 |
| Document understanding | |
| CharXiv RQ | 78.4 |
| CC-OCR | 81.2 |
| OCRBench | 89.4 |
| Spatial intelligence | |
| ERQA | 62.5 |
| CountBench | 97.8 |
| RefCOCO Avg | 92.5 |
| EmbSpatialBench | 84.6 |
| RefSpatialBench | 70.0 |
| Video understanding | |
| VideoMME | 87.7 |
| VideoMMMU | 84.4 |
| MLVU | 86.6 |
| MVBench | 75.5 |
| Visual agent | |
| V* | 94.7 |
| AndroidWorld | 70.3 |
Closed-source reference set
Anchored to the public vendor pricing pages cited under Sources.
| Model | Input $ / 1M tokens | Output $ / 1M tokens | Source |
|---|---|---|---|
| GPT-5 | $1.25 | $10.00 | openai.com/api/pricing |
| Claude Opus 4.5 | $5.00 | $25.00 | docs.anthropic.com |
| Gemini 3.1 Pro | $2.00 | $12.00 | ai.google.dev/gemini-api/docs/pricing |
Open-weight reference set
Open-weight models vary by deployment route; verify the route before adopting.
| Model | Input (typical) | Output (typical) | Source |
|---|---|---|---|
| Qwen3.8-Flash (this model) | $0.16 | $0.47 | QwenCloud via OneInfer (live) |
| GLM-4.5 / 4.6 / 5 | Self-hostable; provider rate varies | Self-hostable; provider rate varies | docs.z.ai/guides/pricing/pricing |
| Qwen2.5 family | Self-hostable; ≈$0.20–$3.00 via 3rd-party | Self-hostable; verify per provider | huggingface.co/Qwen |
| DeepSeek V3.x | Verify DeepSeek pricing | Verify DeepSeek pricing | platform.deepseek.com |
Ready to test the workflow?
Create account & add creditsCost vs control matrix
- Closed APIs: pay-per-token, no GPU ops, retention per vendor terms.
- Open-weight via OneInfer GPU market: full data control, regional residency, no idle capacity if you size to demand.
- Self-host off-OneInfer: maximum control, but you own the serving stack, autoscaling, and incident response.
Workload recommendation
| Workload | How to choose |
|---|---|
| Production agent with steady traffic | Closed API if vendor retention is acceptable; open-weight if residency required. |
| Regulated data (PHI, PCI, classified) | Open-weight on a residency-compliant GPU. |
| Spiky / experimental traffic | Closed API (pay-per-token, no idle GPUs). |
| Cost ceiling unknown | Start closed, mirror to open-weight once the workload stabilizes. |
Frequently asked questions
Can I try Qwen3.8-Flash before integrating it?
Use the OneInfer Qwen3.8-Flash launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.
How should I treat benchmark claims?
Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.
Put Qwen3.8-Flash to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.