Live pricing on OneInfer
Pulled from the OneInfer catalog (OpenRouter and direct provider routes) at request time — rates can change. Capture the timestamp with any benchmark or production decision.
| Model | Provider | Input $/1M tokens | Output $/1M tokens |
|---|---|---|---|
| GLM-5.3 | zai | $140.00 | $440.00 |
| GLM-5.3 | zai | $15.00 | $50.00 |
| GLM-5.3 | zai | $37.00 | $125.00 |
| GLM-5.3 | akashml | $117.00 | $396.00 |
| GLM-5.3 | novita | $15.00 | $50.00 |
| GLM-5.3 | together_ai | $15.00 | $50.00 |
| DeepSeek V4 | deepseek | $30.00 | $120.00 |
| DeepSeek V4 | deepseek | $132.00 | $396.00 |
| DeepSeek V4 | bytedance | $14.00 | $28.00 |
| DeepSeek V4 | bytedance | $174.00 | $348.00 |
| DeepSeek V4 | novita | $14.00 | $28.00 |
| DeepSeek V4 | novita | $160.00 | $320.00 |
| DeepSeek V4 | novita | $30.00 | $120.00 |
Decision snapshot
| Decision factor | GLM-5.3 | DeepSeek V4 Pro 0813 |
|---|---|---|
| Independent intelligence score | 60 | 53 |
| First-party API input / 1M tokens | $1.40 | $0.435 cache miss; $0.003625 cache hit |
| First-party API output / 1M tokens | $4.40 | $0.87 |
| Cost per independent evaluation task | $0.68 | $0.27 |
| Context window | 1M tokens | 1M tokens |
| Maximum output | 128K tokens | Up to 384K tokens |
| Input modalities | Text | Text |
| Weights at verification date | Announced for release | Available under the MIT License |
How we keep this comparison honest
- Pin the exact DeepSeek checkpoint: “DeepSeek V4” here means DeepSeek V4 Pro 0813, not V4 Pro Preview, V4 Flash, or V4 Flash 0731.
- Match max reasoning effort and Artificial Analysis Intelligence Index v4.1.1 for the independent performance table.
- Keep independent results separate from Z.ai's vendor-reported coding and agent benchmark matrix.
- Do not blend price bases: DeepSeek's public endpoint list price and Artificial Analysis' tested exact-checkpoint price snapshot are shown as separate evidence.
- Treat downloadable weights as a deployment option, not an automatic cost win; DeepSeek V4 Pro is a 1.6T-parameter MoE checkpoint with 49B active parameters.
Independent performance and efficiency
Artificial Analysis measured GLM-5.3 max and DeepSeek V4 Pro 0813 max with Intelligence Index v4.1.1. The values below are a dated snapshot of the tested APIs and configurations.
| Artificial Analysis metric | GLM-5.3 (max) | DeepSeek V4 Pro 0813 (max) |
|---|---|---|
| Intelligence Index v4.1.1 | 60 | 53 |
| Cost per Intelligence Index task | $0.68 | $0.27 |
| Output speed | 84.1 tokens/s | 68.3 tokens/s |
| Total evaluation output tokens | 170M | 130M |
What the independent results mean
- GLM-5.3 leads the composite intelligence score by 7 points, the strongest quality signal in this matched evaluation.
- GLM-5.3 generated output about 23% faster on the tested first-party APIs.
- DeepSeek V4 Pro 0813 cost about 60% less per evaluated task and used about 24% fewer output tokens.
- The result favors GLM for measured capability and decode speed, and DeepSeek for evaluated task economics and concision.
Matched coding and agent benchmarks reported by Z.ai
Z.ai reports both exact models in the same GLM-5.3 release matrix. These rows use named shared harnesses, but remain vendor-reported evidence and should be reproduced on a private test set.
| Benchmark | GLM-5.3 | DeepSeek V4 Pro 0813 |
|---|---|---|
| Terminal-Bench 2.1 | 88.2 | 87.9 |
| DeepSWE v1.1 | 66.9 | 62.7 |
| NL2Repo | 58.0 | 61.1 |
| CyberGym | 84.5 | 83.3 |
| Toolathlon Verified | 73.0 | 74.1 |
| AutomationBench v1.0.6 | 48.2 | 43.2 |
| Agents' Last Exam (ALE-CLI) | 28.5 | 25.7 |
| HLE with tools | 62.5 | 60.0 |
Ready to test the workflow?
Create account & add creditsPricing needs two separate views
DeepSeek lists deepseek-v4-pro at $0.003625 per million cached-input tokens, $0.435 per million cache-miss input tokens, and $0.87 per million output tokens. Artificial Analysis reports a different tested exact-checkpoint snapshot of $1.32 input and $3.96 output per million tokens for DeepSeek V4 Pro 0813. Provider route, checkpoint mapping, and observation time can explain material differences, so record the returned model version and live billable rates when testing. GLM-5.3 is listed at $0.26 cached input, $1.40 input, and $4.40 output per million tokens.
Capabilities and API behavior
| Capability | GLM-5.3 | DeepSeek V4 Pro 0813 |
|---|---|---|
| Reasoning modes | Low, high, and max; always enabled | Non-thinking plus low, high, and max reasoning |
| Context window | 1M tokens | 1M tokens |
| Maximum output | 128K tokens | 384K recommended for high/max deployments |
| Text, JSON, and tool workflows | Supported; exact features provider-dependent | JSON output, tool calls, chat prefix, and FIM through the official API |
| Image input | Not supported by the base model | Not supported by V4 Pro 0813 |
| Open-weight availability | Announced; verify release and license | Published under MIT |
| Official serving examples | Verify after weight release | vLLM and SGLang with DSpark speculative decoding |
Self-hosting changes the operational tradeoff
DeepSeek V4 Pro 0813 is already downloadable under MIT and DeepSeek documents vLLM and SGLang deployment, including a 4×GB300 example. The repository is roughly 893 GB and the model has 1.6T total parameters with 49B active, so production serving still requires substantial accelerator memory, interconnect, storage, and operations work. Z.ai had announced but not yet published GLM-5.3 weights at this page's verification point; recheck the official repository before planning GLM self-hosting.
Which model should you choose?
| Workload or priority | Recommended starting point | Why |
|---|---|---|
| Highest independent general capability | GLM-5.3 | It leads the current matched intelligence index by 7 points. |
| Fast hosted generation | GLM-5.3 | It generated about 23% faster in the independent test. |
| Lowest observed task cost | DeepSeek V4 Pro 0813 | Its evaluated task cost was $0.27 versus $0.68. |
| Concise long-running agents | DeepSeek V4 Pro 0813 | It used about 24% fewer evaluation output tokens. |
| Disable reasoning for routine work | DeepSeek V4 Pro 0813 | It supports non-thinking mode; GLM-5.3 reasoning is always enabled. |
| Very large generated outputs | DeepSeek V4 Pro 0813 | DeepSeek documents up to 384K versus GLM's 128K. |
| Open weights available now | DeepSeek V4 Pro 0813 | The MIT-licensed checkpoint and serving recipes are already published. |
| Repository coding and tool agents | Run a private bake-off | Matched vendor rows split across tasks, so no universal coding winner is established. |
Frequently asked questions
Can I try GLM-5.3 before integrating it?
Use the OneInfer GLM-5.3 launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.
How should I treat benchmark claims?
Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.
Which DeepSeek V4 model is compared here?
DeepSeek V4 Pro 0813 at max reasoning effort. It is the official Pro release that superseded the preview. The results do not describe the smaller V4 Flash or V4 Flash 0731 models.
Which model scored higher independently?
Artificial Analysis Intelligence Index v4.1.1 reports GLM-5.3 max at 60 and DeepSeek V4 Pro 0813 max at 53.
Which model is cheaper?
DeepSeek V4 Pro 0813 had the lower independent cost per task, and DeepSeek's public endpoint list rates are lower than GLM-5.3's. Because the independently observed DeepSeek token rates differ from its public list price, verify the exact route, model version, cache behavior, and invoice rate.
Can both models process images?
No. Both the base GLM-5.3 model and DeepSeek V4 Pro 0813 are text-input models.
Can I self-host DeepSeek V4 Pro 0813?
Yes, its MIT-licensed weights and vLLM/SGLang recipes are published. It is still a very large 1.6T-parameter MoE model, and DeepSeek's example uses a 4×GB300 node, so hardware and operating cost require careful sizing.
Put GLM-5.3 to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.