Live pricing on OneInfer
Pulled from the OneInfer catalog at request time — rates can change. Capture the timestamp with any production decision.
| Provider | Input $/1M tokens | Output $/1M tokens | Cache read $/1M tokens |
|---|---|---|---|
| novita | $2.00 | $6.00 | $0.250 |
| together_ai | $2.00 | $6.00 | $0.250 |
Published API pricing per 1M tokens
Source: each provider's own listing, fetched 8 September 2026. One row per provider serving Qwen3.8-2.4T-A95B on OneInfer.
| Provider | Input $/1M | Output $/1M | Cache read $/1M | Batch input / output | Verified at |
|---|---|---|---|---|---|
| Alibaba Qwen (direct) | $2.00 | $6.00 | n/a | $1.00 / $3.00 | 8 Sep 2026 |
| OpenRouter | $2.00 | $6.00 | $0.20 | yes | 8 Sep 2026 |
| DeepInfra | $2.00 | $6.00 | $0.20 | n/a | 8 Sep 2026 |
| Novita | $2.00 | $6.00 | $0.25 | n/a | 8 Sep 2026 |
| Vercel AI Gateway | $2.00 | $6.00 | $0.25 | n/a | 8 Sep 2026 |
| Hugging Face (routes Novita) | $2.00 | $6.00 | n/a | n/a | 8 Sep 2026 |
| Together AI | $2.50 | $6.25 | $0.50 | n/a | 8 Sep 2026 |
Together AI charges 25% more
Together AI is the only tracked provider with a price premium: 25% more on input, 4% more on output, and 2.5× the cache-read rate of OpenRouter and DeepInfra. The premium is consistent with Together AI's other frontier MoE listings, but is the single biggest per-token cost swing across the seven tracked providers.
Alibaba batch cuts cost in half
Alibaba Cloud's batch endpoints run at $1.00 input / $3.00 output per 1M tokens — exactly half the standard rate. For offline reprocessing jobs the batch endpoint is the cheapest published route.
Ready to test the workflow?
Create account & add creditsHow blended cost is calculated
Blended cost assumes a 1:3 input/output token ratio. At 1:1 input:output on the six at-parity providers, the blend is $3.00; at 1:9 it narrows to $2.40. Treat blended as a workload sensitivity test, not a neutral sort key.
Cheapest-first ordering
Alibaba Qwen batch ($1.00 / $3.00) → OpenRouter cache-heavy ($0.20 cache read) → six at-parity providers ($2.00 / $6.00) → Together AI ($2.50 / $6.25, $0.50 cache). Fireworks supports serverless inference but its published rate was not captured at this page's verification.
Frequently asked questions
How much does Qwen3.8-2.4T-A95B cost per token?
$2.00 per 1M input tokens and $6.00 per 1M output tokens as of 8 September 2026 on Alibaba Qwen, OpenRouter, DeepInfra, Novita, Vercel AI Gateway and Hugging Face. Together AI lists $2.50 / $6.25 with a $0.50 cache-read rate. Alibaba's batch endpoint halves cost to $1.00 / $3.00.
Which provider is the cheapest for Qwen3.8-2.4T-A95B?
Alibaba Cloud's batch endpoint at $1.00 input / $3.00 output per 1M tokens. Among real-time endpoints, six providers tie at $2.00 / $6.00.
Where is the live Qwen3.8-2.4T-A95B model page?
The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This page is a focused facet of that entity, not a replacement for it.
How should I treat benchmark or price claims?
Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.
Is Qwen3.8-2.4T-A95B open source?
Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.
Put Qwen3.8-2.4T-A95B to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.