Comparison · most-searched open-weight peer

Qwen3.8-2.4T-A95B vs Kimi K3

Kimi K3 at $4.00 input and $18.00 output per 1M tokens against $2.00 / $6.00 here — Kimi K3 is 2× the input price and 3× the output price. Both are open-weight frontier MoE checkpoints with 1M-token context windows. The buying decision for open-weight frontier buyers is between Kimi K3 and the open-weight Qwen Max (A95B); no shared benchmark table has been published running both under the same harness.

The headline cost gap is 3× on output

Kimi K3 at $4.00 input / $18.00 output per 1M tokens versus Qwen3.8-2.4T-A95B at $2.00 / $6.00. On a 1:3 input:output workload the blended rate is $14.50 on Kimi K3 against $5.00 on A95B — A95B is 65% cheaper on the same workload. On a 1:9 long-output workload the gap widens to 75%.

Workload (input:output)Qwen3.8-2.4T-A95BKimi K3Saving on A95B
1:1 (chat-shaped)$4.00 / 1M blended$11.00 / 1M blended64%
1:3 (agent-shaped)$5.00 / 1M blended$14.50 / 1M blended65%
1:9 (long-output)$2.40 / 1M blended$9.80 / 1M blended75%
Input only (rerank / classify)$2.00 / 1M$4.00 / 1M50%
Output only (synthesis)$6.00 / 1M$18.00 / 1M67%

Both are open-weight frontier MoE checkpoints

Qwen3.8-2.4T-A95B: 2.4T total parameters, 95B active per token, 262K native context extensible to 1M via YaRN, qwen3.8-max license. Kimi K3: published by Moonshot AI with a 1M-token context window, Moonshot's published Kimi K3 license. Both route through OneInfer via OpenAI-compatible endpoints. Both ship open weights.

No published same-harness benchmark table

Neither Moonshot nor Alibaba publishes a shared benchmark table running Kimi K3 and A95B on the same harness. A95B scores 86.6 on Terminal Bench 2.1, 92.6 on GPQA Diamond, 67.7 on SWE-bench Pro, 73.5 on FrontierSWE — all on Qwen's published harnesses. Kimi K3's published scores are on Moonshot's own eval harness. State the absence before quoting any side-by-side number.

When to pick Kimi K3

  • You specifically want Moonshot's published Kimi K3 license terms over the qwen3.8-max license.
  • You have a workload that has been pre-tuned on Kimi K3 and a switch is more costly than the 3× output premium.
  • You are running inference on hardware or a serving stack that has been benchmarked against Kimi K3 already.

Ready to test the workflow?

Create account & add credits

When to pick A95B

  • Per-token cost dominates the buying decision. Output-heavy workloads in particular.
  • You are running workloads where 262K native / 1M extended context is required and YaRN extension is acceptable.
  • You want the Qwen vendor benchmark record reproduced on Anthropic / OpenAI / Google frontier harnesses.

Caveats

  • Kimi K3's published license is Moonshot's Kimi K3 license, not Apache or MIT. Read the license tag on the official Hugging Face repository before self-hosting at scale.
  • Cache-read rates for Kimi K3 are not consistently published; the headline rates above are uncached.
  • A95B requires thinking mode on every call. Kimi K3 has not been documented as requiring thinking mode.

Frequently asked questions

Is Kimi K3 more expensive than Qwen3.8-2.4T-A95B?

Yes — Kimi K3 is 2× the input price and 3× the output price: $4.00 / $18.00 against $2.00 / $6.00 per 1M tokens. On a 1:3 input:output workload the blended rate is $14.50 against $5.00, a 65% saving on A95B.

Are Qwen3.8-2.4T-A95B and Kimi K3 the same kind of model?

Both are open-weight frontier MoE checkpoints with 1M-token context windows. They differ on parameter count, active count per token, and the published license terms. No shared benchmark table has been published running both under the same harness.

Where is the canonical Qwen3.8-2.4T-A95B model page?

The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This comparison is a focused facet of that entity, not a replacement for it.

How should I treat benchmark or price claims on this page?

Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.

Is Qwen3.8-2.4T-A95B open source?

Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.

Put Qwen3.8-2.4T-A95B to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.