vs Qwen3.8-Max — disambiguation, not a competition
Qwen3.8-Max is the hosted variant of the same weights. Max adds vision input, non-thinking mode, 1M default context, and built-in tools. This checkpoint is text-only and cannot disable thinking. If you specifically need vision or non-thinking mode, Max is the answer. If you specifically need the open weights or the lowest per-token rate, A95B is the answer.
| Dimension | Qwen3.8-2.4T-A95B | Qwen3.8-Max | Full comparison |
|---|---|---|---|
| Weights | Open (qwen3.8-max license) | Hosted, closed | Compare → |
| Vision input | No (text-only) | Yes | |
| Non-thinking mode | No | Yes | |
| Default context | 262K native, 1M extended | 1M | |
| Input / Output per 1M | $2.00 / $6.00 | Hosted, varies |
vs Kimi K3 — most-searched genuine competitor
Kimi K3 at $4.00 input / $18.00 output per 1M tokens versus $2.00 / $6.00 here — 3× the output cost. Both are open-weight frontier MoE checkpoints. This is the real buying decision for open-weight buyers.
Choosing by workload
Lowest per-token cost → DeepSeek V4 Pro. Cheapest input + output together → GLM 5.2. Vision input or non-thinking mode → Qwen3.8-Max. Hardest open-weight peer at the same frontier tier → Kimi K3. Hardest reasoning depth, willing to pay more → Qwen3.8-Max (hosted) or Fable 5 / Opus 4.8 (closed).
Ready to test the workflow?
Create account & add creditsRejected alternatives, with reasons
- vs Gemma 4 31B — 31B against 2.4T. Fails the adjacent-capability-tier gate; the comparison is not meaningful and the page reads as filler.
- vs GPT-5.6 Sol — Qwen benchmarks against it, but there are no open weights and no OneInfer routing path. Cannot serve the query that lands here.
- vs Qwen3-Coder — Different generation and specialisation; cannibalises the Qwen family hub.
Frequently asked questions
What is the closest open-weight alternative to Qwen3.8-2.4T-A95B?
Kimi K3, at $4.00 / $18.00 per 1M tokens — 3× the output cost. For cheaper per-token pricing, DeepSeek V4 Pro at $1.80 / $3.60 and GLM 5.2 at $1.40 / $4.40 both cost less.
What is the difference between Qwen3.8-2.4T-A95B and Qwen3.8-Max?
Same weights. Qwen3.8-Max is the hosted variant: adds vision input, non-thinking mode, 1M default context, and built-in tools. Qwen3.8-2.4T-A95B is text-only, requires thinking mode, and ships as open weights.
Where is the live Qwen3.8-2.4T-A95B model page?
The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This page is a focused facet of that entity, not a replacement for it.
How should I treat benchmark or price claims?
Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.
Is Qwen3.8-2.4T-A95B open source?
Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.
Put Qwen3.8-2.4T-A95B to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.