Model facet · Providers

Qwen3.8-Flash inference providers

Qwen3.8-Flash is a hosted Flash API; this hub tracks every tracked inference route once providers publish their pricing and routing — expect the snapshot to fill in as providers onboard Qwen3.8-Flash.

Tracked providers

ProviderInput /1M tokensOutput /1M tokensProvenance
openrouter$0.15$0.47vendor reported

What is Qwen3.8-Flash?

Qwen3.8-Flash is Alibaba's multimodal flash-tier model, released 26 August 2026 and priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens. It accepts a 1,000,000-token context with image and text input, and returns up to 131,072 tokens. Comparable flash-tier models range from $0.03 to $0.16 on input.

Providers

Each row will be labeled by provenance: independently verified listings versus a provider's own vendor-reported number. Routing, quantization, and uptime can all change — record the provider and date alongside any benchmark or latency claim you rely on.

Ready to test the workflow?

Create account & add credits

When Qwen3.8-Flash is not the right choice

If the workload is the hardest desktop automation (OSWorld 2.0 binary ≈ 19.4) or scripted RPA, Flash trails Claude Opus 4.6 and DeepSeek V4 Flash 0731 — pick those instead. Qwen3.8-Flash is also not the cheapest option (5.3× DeepSeek V4 Flash 0731 on input) and not the strongest on long-horizon reasoning; Qwen3.8-Max exists for the latter.

Frequently asked questions

Where is the live Qwen3.8-Flash model page?

The canonical model page with current OneInfer pricing, capabilities, and availability is /models/Qwen/Qwen3.8-Flash. This page is a focused facet of that entity, not a replacement for it.

How should I treat benchmark or price claims?

Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.

Put Qwen3.8-Flash to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.