Model facet · Providers

7 Providers Serving Qwen3.8-2.4T-A95B, Cheapest First

Seven providers currently publish Qwen3.8-2.4T-A95B pricing. Six charge $2.00 / $6.00 per 1M tokens; the live catalog shows Together AI at the same $2.00 / $6.00. Cache read at $0.25 per 1M tokens.

Live pricing on OneInfer

Pulled from the OneInfer catalog at request time — rates can change. Capture the timestamp with any production decision.

ProviderInput $/1M tokensOutput $/1M tokensCache read $/1M tokens
novita$2.00$6.00$0.250
together_ai$2.00$6.00$0.250

How to read this table

Each row is labeled by provenance: independently verified listings versus a provider's own vendor-reported number. Routing, quantization, and uptime can all change — record the provider and date alongside any benchmark or latency claim you rely on.

Provider matrix

ProviderInput $/1MOutput $/1MCache readBatchProvenance
Alibaba Qwen (direct)$2.00$6.00n/a$1.00 / $3.00vendor_reported
OpenRouter$2.00$6.00$0.20yesverified
DeepInfra$2.00$6.00$0.20n/averified
Novita$2.00$6.00$0.25n/averified
Vercel AI Gateway$2.00$6.00$0.25n/averified
Hugging Face$2.00$6.00n/an/averified
Together AI$2.50$6.25$0.50n/averified

Ready to test the workflow?

Create account & add credits

Fireworks supports serverless inference

Fireworks lists Qwen3.8-Max for serverless inference and notes fine-tuning is not supported on that listing. A separate, captured rate for Qwen3.8-2.4T-A95B on Fireworks was not in scope at this page's verification — check the live listing before quoting per-token cost.

Routing notes

Hugging Face's hosted route is backed by Novita in this case. Vercel AI Gateway exposes the model under the OpenAI-compatible path used by most existing SDKs. Each provider may run at a different precision (FP8 vs BF16) and a different concurrency ceiling; verify against the live provider dashboard before serving production traffic.

Frequently asked questions

Where can I run Qwen3.8-2.4T-A95B?

Seven providers publish pricing as of 8 September 2026: Alibaba Qwen (direct), OpenRouter, DeepInfra, Novita, Vercel AI Gateway, Hugging Face, and Together AI. Fireworks lists serverless support; its per-token rate was not captured at this page's verification.

Why does Together AI cost more?

Together AI's standard rate is 25% above the six at-parity providers on input and 4% above on output. Its cache-read rate is 2.5× OpenRouter and DeepInfra. The premium is consistent with Together's other frontier MoE listings.

Where is the live Qwen3.8-2.4T-A95B model page?

The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This page is a focused facet of that entity, not a replacement for it.

How should I treat benchmark or price claims?

Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.

Is Qwen3.8-2.4T-A95B open source?

Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.

Put Qwen3.8-2.4T-A95B to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.