Live pricing on OneInfer
Pulled from the OneInfer catalog (OpenRouter and direct provider routes) at request time — rates can change. Capture the timestamp with any benchmark or production decision.
| Model | Provider | Input $/1M tokens | Output $/1M tokens |
|---|---|---|---|
| GLM-5.3 | zai | $140.00 | $440.00 |
| GLM-5.3 | zai | $15.00 | $50.00 |
| GLM-5.3 | zai | $37.00 | $125.00 |
| GLM-5.3 | akashml | $117.00 | $396.00 |
| GLM-5.3 | novita | $15.00 | $50.00 |
| GLM-5.3 | together_ai | $15.00 | $50.00 |
| Kimi K3 | novita | $300.00 | $1500.00 |
Decision snapshot
| Decision factor | GLM-5.3 | Kimi K3 |
|---|---|---|
| Independent intelligence score | 60 | 60 |
| Observed input price / 1M tokens | $1.40 first-party | $3.00 in tested API |
| Observed output price / 1M tokens | $4.40 first-party | $15.00 in tested API |
| Cost per independent evaluation task | $0.68 | $0.84 |
| Context window | 1M tokens | 1M tokens |
| Maximum output | 128K tokens | Provider/runtime-dependent; verify |
| Input modalities | Text | Text, image, and video |
| Weights at verification date | Announced for August 28 release | Available under the Kimi K3 License |
How we keep this comparison honest
- Match the evaluated variants: the independent table compares GLM-5.3 max with Kimi K3 max on Artificial Analysis Intelligence Index v4.1.1.
- Keep independent measurements separate from Z.ai's vendor-reported coding scores, even when the model names appear in the same table.
- Label the price basis: GLM figures are Z.ai first-party list prices; Kimi figures are prices observed for the API tested by Artificial Analysis and may vary by provider.
- Treat open weights as deployment control, not free inference: Kimi K3 has 2.8T total parameters, 104B active parameters, and a very large weight distribution.
- Compare base-model modalities explicitly: GLM-5.3 is text-only, while Moonshot documents native image and video understanding for Kimi K3.
Independent performance and efficiency
Artificial Analysis measured both max variants with Intelligence Index v4.1.1. These numbers describe the tested APIs and configurations, not every provider route or reasoning setting.
| Artificial Analysis metric | GLM-5.3 (max) | Kimi K3 (max) |
|---|---|---|
| Intelligence Index v4.1.1 | 60 | 60 |
| Cost per Intelligence Index task | $0.68 | $0.84 |
| Output speed | 84.1 tokens/s | 34.3 tokens/s |
| Total evaluation output tokens | 170M | 130M |
What the independent results mean
- Measured general intelligence is a tie, so the composite score alone does not select a winner.
- GLM-5.3 cost about 19% less per evaluated task and generated output roughly 2.5× faster.
- Kimi K3 used about 24% fewer output tokens, an efficiency advantage that can matter for long agent runs.
- The speed and cost results are provider-specific observations; benchmark your actual route, tool policy, and latency region before production selection.
Ready to test the workflow?
Create account & add creditsMatched coding benchmarks reported by Z.ai
Z.ai publishes both models in the same GLM-5.3 release table. These are useful matched rows, but they remain vendor-reported evidence and should be reproduced independently before procurement decisions.
| Benchmark | GLM-5.3 | Kimi K3 |
|---|---|---|
| Terminal-Bench 2.1 | 88.2 | 88.3 |
| Terminal-Bench 3.0 | 28.3 | 17.4 |
| DeepSWE v1.1 | 66.9 | 67.5 |
Capabilities and deployment control
| Capability | GLM-5.3 | Kimi K3 |
|---|---|---|
| Reasoning variants | Low, high, and max; reasoning always enabled | Max and low variants independently tracked; verify provider controls |
| Multimodal understanding | Text only | Native text, image, and video |
| Context window | 1M tokens | 1,048,576 tokens |
| Open-weight availability | Announced, not yet published at verification | Weights published on Hugging Face |
| Official serving examples | Verify after weight release | vLLM and SGLang |
| Model scale | Verify from released model card | 2.8T total; 104B active parameters |
Open weights do not mean lightweight deployment
Kimi K3 weights are available under the Kimi K3 License, and Moonshot provides vLLM and SGLang examples. The model is nevertheless enormous: 2.8 trillion total parameters with 104 billion active per token, and its published repository is roughly 1.56 TB. Review the license, precision, memory, networking, and multi-GPU serving cost before choosing self-hosting. Z.ai announced GLM-5.3 weights for August 28, 2026; they were not yet available at this page's August 27 verification point.
Which model should you choose?
| Workload or priority | Recommended starting point | Why |
|---|---|---|
| Highest measured general intelligence | Test both | The max variants tie on the current independent index. |
| Fast hosted text and coding agents | GLM-5.3 | It generated about 2.5× faster in the independent test. |
| Lower observed API task cost | GLM-5.3 | Its evaluated task cost was $0.68 versus $0.84. |
| Image or video understanding | Kimi K3 | Kimi is natively multimodal; base GLM-5.3 is text-only. |
| Concise agent output | Kimi K3 | It used about 24% fewer evaluation output tokens. |
| Open weights available today | Kimi K3 | Its weights and serving examples are already published. |
| Self-hosting on limited hardware | Neither by default | Kimi is exceptionally large, and GLM weights were not yet available to validate. |
| Terminal or repository coding | Run a private bake-off | The matched vendor table is split across benchmarks and does not establish one universal winner. |
Frequently asked questions
Can I try GLM-5.3 before integrating it?
Use the OneInfer GLM-5.3 launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.
How should I treat benchmark claims?
Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.
Which model scores higher independently?
Neither in the current max-effort comparison. Artificial Analysis Intelligence Index v4.1.1 reports both GLM-5.3 max and Kimi K3 max at 60.
Is GLM-5.3 cheaper than Kimi K3?
In the tested API snapshot, yes: GLM-5.3 cost $0.68 per independent evaluation task versus $0.84 for Kimi K3 and had lower observed input and output token rates. Provider routing, caching, and token use can change the result.
Can both models process images and video?
No. Moonshot documents native image and video understanding for Kimi K3. The base GLM-5.3 model accepts text only.
Can I self-host Kimi K3?
Its weights are published with vLLM and SGLang examples, but self-hosting is a major infrastructure project because the model has 2.8T total and 104B active parameters. Review the Kimi K3 License and size the multi-GPU system before deployment.
Are GLM-5.3 weights available?
At this page's August 27, 2026 verification point, Z.ai had announced the weights for August 28 but had not yet published them. Recheck the official model repository after the announced date.
Put GLM-5.3 to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.