Live pricing on OneInfer
Pulled from the OneInfer catalog at request time — rates can change. Capture the timestamp with any production decision.
| Provider | Input $/1M tokens | Output $/1M tokens | Cache read $/1M tokens |
|---|---|---|---|
| novita | $2.00 | $6.00 | $0.250 |
| together_ai | $2.00 | $6.00 | $0.250 |
How to read this table
Each row is labeled by provenance: independently verified listings versus a provider's own vendor-reported number. Routing, quantization, and uptime can all change — record the provider and date alongside any benchmark or latency claim you rely on.
Provider matrix
| Provider | Input $/1M | Output $/1M | Cache read | Batch | Provenance |
|---|---|---|---|---|---|
| Alibaba Qwen (direct) | $2.00 | $6.00 | n/a | $1.00 / $3.00 | vendor_reported |
| OpenRouter | $2.00 | $6.00 | $0.20 | yes | verified |
| DeepInfra | $2.00 | $6.00 | $0.20 | n/a | verified |
| Novita | $2.00 | $6.00 | $0.25 | n/a | verified |
| Vercel AI Gateway | $2.00 | $6.00 | $0.25 | n/a | verified |
| Hugging Face | $2.00 | $6.00 | n/a | n/a | verified |
| Together AI | $2.50 | $6.25 | $0.50 | n/a | verified |
Ready to test the workflow?
Create account & add creditsFireworks supports serverless inference
Fireworks lists Qwen3.8-Max for serverless inference and notes fine-tuning is not supported on that listing. A separate, captured rate for Qwen3.8-2.4T-A95B on Fireworks was not in scope at this page's verification — check the live listing before quoting per-token cost.
Routing notes
Hugging Face's hosted route is backed by Novita in this case. Vercel AI Gateway exposes the model under the OpenAI-compatible path used by most existing SDKs. Each provider may run at a different precision (FP8 vs BF16) and a different concurrency ceiling; verify against the live provider dashboard before serving production traffic.
Frequently asked questions
Where can I run Qwen3.8-2.4T-A95B?
Seven providers publish pricing as of 8 September 2026: Alibaba Qwen (direct), OpenRouter, DeepInfra, Novita, Vercel AI Gateway, Hugging Face, and Together AI. Fireworks lists serverless support; its per-token rate was not captured at this page's verification.
Why does Together AI cost more?
Together AI's standard rate is 25% above the six at-parity providers on input and 4% above on output. Its cache-read rate is 2.5× OpenRouter and DeepInfra. The premium is consistent with Together's other frontier MoE listings.
Where is the live Qwen3.8-2.4T-A95B model page?
The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This page is a focused facet of that entity, not a replacement for it.
How should I treat benchmark or price claims?
Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.
Is Qwen3.8-2.4T-A95B open source?
Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.
Put Qwen3.8-2.4T-A95B to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.