Live pricing on OneInfer
Pulled from the OneInfer catalog (OpenRouter and direct provider routes) at request time — rates can change. Capture the timestamp with any benchmark or production decision.
| Model | Provider | Input $/1M tokens | Output $/1M tokens |
|---|---|---|---|
| GLM-5.3 | zai | $140.00 | $440.00 |
| GLM-5.3 | zai | $15.00 | $50.00 |
| GLM-5.3 | zai | $37.00 | $125.00 |
| GLM-5.3 | akashml | $117.00 | $396.00 |
| GLM-5.3 | novita | $15.00 | $50.00 |
| GLM-5.3 | together_ai | $15.00 | $50.00 |
| Qwen | novita | $38.00 | $40.00 |
| Qwen | novita | $9.00 | $58.00 |
| Qwen | novita | $38.00 | $155.00 |
| Qwen | novita | $15.00 | $150.00 |
| Qwen | cerebras | $99.00 | $149.00 |
| Qwen | novita | $30.00 | $300.00 |
| Qwen | novita | $40.00 | $320.00 |
| Qwen | novita | $30.00 | $240.00 |
| Qwen | novita | $25.00 | $200.00 |
| Qwen | novita | $60.00 | $360.00 |
| Qwen | thinkingmachines | $400.00 | $1000.00 |
| Qwen | novita | $60.00 | $360.00 |
| Qwen | thinkingmachines | $186.00 | $559.50 |
| Qwen | groq | $60.00 | $300.00 |
| Qwen | novita | $24.80 | $148.50 |
| Qwen | thinkingmachines | $54.00 | $133.50 |
| Qwen | akashml | $10.00 | $90.00 |
| Qwen | novita | $125.00 | $375.00 |
| Qwen | novita | $200.00 | $600.00 |
| Qwen | together_ai | $200.00 | $600.00 |
| Qwen | akashml | $22.50 | $198.00 |
| Qwen | novita | $42.00 | $300.00 |
| Qwen | novita | $15.00 | $47.00 |
| Qwen | novita | $200.00 | $600.00 |
| Qwen | openrouter | $15.00 | $47.00 |
| Qwen | openrouter | $15.00 | $47.00 |
| Qwen | novita | — | — |
| Qwen | openrouter | — | — |
| Qwen | openrouter | — | — |
| Qwen | openrouter | — | — |
Decision snapshot
| Decision factor | GLM-5.3 | Qwen3.8-Max |
|---|---|---|
| API model ID | glm-5.3 | qwen3.8-max |
| Independent intelligence score | 60 | 58 |
| Input price / 1M tokens | $1.40 | $2.00 |
| Cached input / 1M tokens | $0.26 | $0.25 implicit; $0.17 explicit read |
| Output price / 1M tokens | $4.40 | $6.00 |
| Context window | 1M tokens | 1M tokens |
| Maximum output | 128K tokens | 131K tokens |
| Input modalities | Text | Text, image, and video |
| Managed endpoint weights | Weights announced for release | Max service is proprietary; related base weights available |
How we keep this comparison honest
- Pin the rival to qwen3.8-max. Results for Qwen3.8 27B, Flash, older Qwen-Max releases, or community routes are not substituted.
- Use Artificial Analysis Intelligence Index v4.1.1 for both models in the independent table.
- Keep Z.ai's same-table coding and agent results labelled as vendor-reported evidence.
- Separate the managed Qwen3.8-Max endpoint from Qwen3.8-2.4T-A95B weights: Qwen documents Max-only additions such as vision, non-thinking mode, one-million-token defaults, and built-in tools.
- Record cache mode and region when pricing Qwen because implicit cache, explicit cache creation, and explicit cache reads have different rates.
Independent performance and efficiency
Artificial Analysis measured both reasoning models with Intelligence Index v4.1.1 on their first-party APIs. These figures are a dated snapshot and should be reproduced with your prompts and deployment region.
| Artificial Analysis metric | GLM-5.3 (max) | Qwen3.8-Max |
|---|---|---|
| Intelligence Index v4.1.1 | 60 | 58 |
| Cost per Intelligence Index task | $0.68 | $0.91 |
| Output speed | 84.1 tokens/s | 20.9 tokens/s |
| Total evaluation output tokens | 170M | 150M |
What the independent results mean
- GLM-5.3 leads measured general intelligence by 2 points, so the quality gap is narrow enough to justify a workload-specific bake-off.
- GLM generated output about four times faster, the largest practical separation in this comparison.
- GLM cost about 25% less per evaluated task despite Qwen using about 12% fewer output tokens.
- Qwen remains competitive on intelligence while adding visual input and non-thinking operation that base GLM-5.3 does not offer.
Matched coding and agent benchmarks reported by Z.ai
Z.ai lists GLM-5.3 and Qwen3.8-Max in the same release matrix. The rows are useful for task-level direction, but they remain vendor-reported and do not establish production reliability.
| Benchmark | GLM-5.3 | Qwen3.8-Max |
|---|---|---|
| Terminal-Bench 2.1 | 88.2 | 86.6 |
| DeepSWE v1.1 | 66.9 | 56.6 |
| NL2Repo | 58.0 | 55.9 |
| CyberGym | 84.5 | 78.5 |
| Toolathlon Verified | 73.0 | 72.5 |
| AutomationBench v1.0.6 | 48.2 | 39.8 |
| Agents' Last Exam (ALE-CLI) | 28.5 | 27.0 |
| HLE with tools | 62.5 | 56.2 |
Ready to test the workflow?
Create account & add creditsCapabilities and managed tools
| Capability | GLM-5.3 | Qwen3.8-Max |
|---|---|---|
| Reasoning control | Low, high, and max; always enabled | Thinking or non-thinking mode |
| Multimodal understanding | Text only | Text, image, and video |
| Function calling and structured output | Supported | Supported |
| Built-in hosted tools | Provider-dependent | Code interpreter, web extractor, web search, text-to-image search, and image-to-image search |
| Context and output | 1M context; 128K output | 1M context; 131K output; up to 262K reasoning |
| Prefix completion | Verify provider support | Supported through Partial Mode |
| Fine-tuning | Verify provider support | Qwen Cloud documents support |
Qwen caching and regional pricing
Qwen Cloud lists $2 input and $6 output per million tokens for qwen3.8-max, with implicit cached input at $0.25, explicit cache creation at $2.50, and explicit cache reads at $0.17. Alibaba Model Studio also publishes region-specific CNY rates. Calculate costs using the deployment region, cache strategy, and expected prompt reuse rather than treating one cache number as universal.
Managed Max is not identical to the downloadable checkpoint
Qwen3.8-Max is based on the downloadable Qwen3.8-2.4T-A95B checkpoint, a 2.4T-parameter MoE with 95B active parameters. Qwen explicitly identifies vision input, non-thinking support, a one-million-token default, and official built-in tools as additions to the managed Max service. Self-hosting the related checkpoint therefore does not guarantee endpoint parity, and its enormous size requires substantial distributed serving infrastructure. Review the Qwen3.8 license and reproduce the exact features you need before deployment.
Which model should you choose?
| Workload or priority | Recommended starting point | Why |
|---|---|---|
| Highest independent intelligence score | GLM-5.3 | It leads by 2 points in the current matched evaluation. |
| Fast interactive coding agents | GLM-5.3 | It generated output about four times faster. |
| Lower measured task cost | GLM-5.3 | It cost $0.68 versus $0.91 per independent evaluation task. |
| Image or video understanding | Qwen3.8-Max | Qwen accepts both; base GLM-5.3 is text-only. |
| Routine tasks without reasoning | Qwen3.8-Max | Qwen can disable thinking; GLM reasoning is always enabled. |
| Managed search and code tools | Qwen3.8-Max | Its endpoint documents a broad built-in tool catalog. |
| Downloadable related weights | Qwen3.8 base checkpoint | Weights are available, but do not assume full Max endpoint parity. |
| Repository coding and long-horizon agents | Test both | GLM leads most matched rows, but private task success and latency should decide. |
Frequently asked questions
Can I try GLM-5.3 before integrating it?
Use the OneInfer GLM-5.3 launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.
How should I treat benchmark claims?
Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.
Which Qwen model is compared here?
The managed qwen3.8-max endpoint. The benchmark and pricing data do not describe Qwen3.8 27B, Flash, older Qwen-Max releases, or an arbitrary self-hosted checkpoint.
Which model scored higher independently?
Artificial Analysis Intelligence Index v4.1.1 reports GLM-5.3 max at 60 and Qwen3.8-Max at 58.
Which model is faster and cheaper?
In the independent snapshot, GLM-5.3 generated at 84.1 versus 20.9 tokens per second and cost $0.68 versus $0.91 per task. Cache behavior, region, reasoning mode, and token use can change real application cost.
Can both models process images and video?
No. Qwen3.8-Max accepts text, image, and video input. The base GLM-5.3 model accepts text only.
Can I self-host Qwen3.8-Max?
Qwen publishes the related Qwen3.8-2.4T-A95B weights, but states that the managed Max endpoint adds capabilities including vision, non-thinking mode, one-million-token defaults, and built-in tools. Treat the checkpoint as related rather than feature-identical, review its license, and account for its 2.4T/95B-active scale.
Put GLM-5.3 to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.