Model facet · Pricing

Qwen3.8-2.4T-A95B Pricing: $2.00 to $2.50 per 1M Input

Qwen3.8-2.4T-A95B lists at $2.00 input and $6.00 output per 1M tokens as of 8 September 2026. Together AI is the only tracked provider charging more, at $2.50 / $6.25 with a $0.50 cache-read rate. Alibaba Qwen's batch discounts halve the standard rate.

Live pricing on OneInfer

Pulled from the OneInfer catalog at request time — rates can change. Capture the timestamp with any production decision.

ProviderInput $/1M tokensOutput $/1M tokensCache read $/1M tokens
novita$2.00$6.00$0.250
together_ai$2.00$6.00$0.250

Published API pricing per 1M tokens

Source: each provider's own listing, fetched 8 September 2026. One row per provider serving Qwen3.8-2.4T-A95B on OneInfer.

ProviderInput $/1MOutput $/1MCache read $/1MBatch input / outputVerified at
Alibaba Qwen (direct)$2.00$6.00n/a$1.00 / $3.008 Sep 2026
OpenRouter$2.00$6.00$0.20yes8 Sep 2026
DeepInfra$2.00$6.00$0.20n/a8 Sep 2026
Novita$2.00$6.00$0.25n/a8 Sep 2026
Vercel AI Gateway$2.00$6.00$0.25n/a8 Sep 2026
Hugging Face (routes Novita)$2.00$6.00n/an/a8 Sep 2026
Together AI$2.50$6.25$0.50n/a8 Sep 2026

Together AI charges 25% more

Together AI is the only tracked provider with a price premium: 25% more on input, 4% more on output, and 2.5× the cache-read rate of OpenRouter and DeepInfra. The premium is consistent with Together AI's other frontier MoE listings, but is the single biggest per-token cost swing across the seven tracked providers.

Alibaba batch cuts cost in half

Alibaba Cloud's batch endpoints run at $1.00 input / $3.00 output per 1M tokens — exactly half the standard rate. For offline reprocessing jobs the batch endpoint is the cheapest published route.

Ready to test the workflow?

Create account & add credits

How blended cost is calculated

Blended cost assumes a 1:3 input/output token ratio. At 1:1 input:output on the six at-parity providers, the blend is $3.00; at 1:9 it narrows to $2.40. Treat blended as a workload sensitivity test, not a neutral sort key.

Cheapest-first ordering

Alibaba Qwen batch ($1.00 / $3.00) → OpenRouter cache-heavy ($0.20 cache read) → six at-parity providers ($2.00 / $6.00) → Together AI ($2.50 / $6.25, $0.50 cache). Fireworks supports serverless inference but its published rate was not captured at this page's verification.

Frequently asked questions

How much does Qwen3.8-2.4T-A95B cost per token?

$2.00 per 1M input tokens and $6.00 per 1M output tokens as of 8 September 2026 on Alibaba Qwen, OpenRouter, DeepInfra, Novita, Vercel AI Gateway and Hugging Face. Together AI lists $2.50 / $6.25 with a $0.50 cache-read rate. Alibaba's batch endpoint halves cost to $1.00 / $3.00.

Which provider is the cheapest for Qwen3.8-2.4T-A95B?

Alibaba Cloud's batch endpoint at $1.00 input / $3.00 output per 1M tokens. Among real-time endpoints, six providers tie at $2.00 / $6.00.

Where is the live Qwen3.8-2.4T-A95B model page?

The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-2.4T-A95B. This page is a focused facet of that entity, not a replacement for it.

How should I treat benchmark or price claims?

Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.

Is Qwen3.8-2.4T-A95B open source?

Yes — weights are released under the qwen3.8-max license. Confirm the license tag against the official Hugging Face repository before self-hosting at scale.

Put Qwen3.8-2.4T-A95B to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.