Live pricing on OneInfer
Pulled from the OneInfer catalog (OpenRouter and direct provider routes) at request time — rates can change. Capture the timestamp with any benchmark or production decision.
| Model | Provider | Input $/1M tokens | Output $/1M tokens |
|---|---|---|---|
| GLM-5.3 | zai | $140.00 | $440.00 |
| GLM-5.3 | zai | $15.00 | $50.00 |
| GLM-5.3 | zai | $37.00 | $125.00 |
| GLM-5.3 | akashml | $117.00 | $396.00 |
| GLM-5.3 | novita | $15.00 | $50.00 |
| GLM-5.3 | together_ai | $15.00 | $50.00 |
| Gemini 3.1 Pro Preview | $200.00 | $1200.00 | |
| Gemini 3.1 Pro Preview | openrouter | $200.00 | $1200.00 |
Decision snapshot
| Decision factor | GLM-5.3 | Gemini 3.1 Pro Preview |
|---|---|---|
| API model ID | glm-5.3 | gemini-3.1-pro-preview |
| Independent intelligence score | 60 | 48 |
| Input price / 1M tokens | $1.40 | $2.00 up to 200K prompt; $4.00 above |
| Cached input / 1M tokens | $0.26 | $0.20 up to 200K prompt; $0.40 above |
| Output price / 1M tokens | $4.40 | $12.00 up to 200K prompt; $18.00 above |
| Context window | 1M tokens | 1M tokens |
| Maximum output | 128K tokens | 64K tokens |
| Input modalities | Text | Text, image, audio, and video |
| Lifecycle and weights | API; weights announced for release | Preview API; proprietary weights |
How we keep this comparison honest
- Use the exact endpoint name: Google currently documents gemini-3.1-pro-preview, not a generally available Gemini 3.1 Pro production endpoint.
- Use one independent harness: the main performance table comes from Artificial Analysis Intelligence Index v4.1.1 for both models.
- Do not hide reasoning configuration: GLM-5.3 is shown at max effort; Gemini 3.1 Pro uses dynamic thinking and defaults to high unless configured otherwise.
- Compare task economics as well as list prices: Gemini has higher headline output pricing but used 56M evaluation output tokens versus GLM-5.3's 170M.
- Apply Gemini's prompt-size tier: its input, cached-input, and output rates all increase when the prompt exceeds 200K tokens.
Independent performance and efficiency
Artificial Analysis measured both models with Intelligence Index v4.1.1. The figures below describe the tested APIs and configurations, not every possible reasoning level or provider route.
| Artificial Analysis metric | GLM-5.3 (max) | Gemini 3.1 Pro Preview |
|---|---|---|
| Intelligence Index v4.1.1 | 60 | 48 |
| Cost per Intelligence Index task | $0.68 | $0.33 |
| Output speed | 84.1 tokens/s | 117.4 tokens/s |
| Total evaluation output tokens | 170M | 56M |
What the independent results mean
- GLM-5.3 leads the composite intelligence score by 12 points, the clearest quality separation among the tested metrics.
- Gemini 3.1 Pro Preview cost about 51% less per evaluated task despite its higher standard output-token price.
- Gemini generated output about 40% faster and used about 67% fewer output tokens in the evaluation.
- The result favors GLM for measured text reasoning and Gemini for speed, concision, multimodal input, and evaluated task cost.
Ready to test the workflow?
Create account & add creditsCapabilities and API behavior
| Capability | GLM-5.3 | Gemini 3.1 Pro Preview |
|---|---|---|
| Reasoning control | Low, high, and max; always enabled | Low, medium, and high; high dynamic default |
| Multimodal understanding | Text only | Text, image, audio, video, and PDFs |
| Function calling and structured output | Supported | Supported, including use with built-in tools |
| Built-in tools | Provider-dependent | Google Search, Maps, URL context, file search, code execution, and computer use |
| Context caching | Supported | Supported with token and hourly storage charges |
| Knowledge cutoff | Verify current vendor disclosure | January 2025 |
Long-context pricing changes the winner
Gemini 3.1 Pro Preview doubles input pricing from $2 to $4 per million tokens when the prompt exceeds 200K tokens, raises output pricing from $12 to $18, and raises cached input from $0.20 to $0.40. Cached context also costs $4.50 per million tokens per hour to store. Model long-document and repository-agent costs at realistic prompt sizes instead of using the below-200K headline rate.
Preview and deployment tradeoffs
Google labels all Gemini 3 models, including Gemini 3.1 Pro, as preview. Preview endpoints can change before general availability, so pin the exact model ID and maintain regression tests. Gemini weights are proprietary. Z.ai announced a GLM-5.3 weight release for August 28, 2026; at this page's August 27 verification point, the weights were not yet published.
Which model should you choose?
| Workload or priority | Recommended starting point | Why |
|---|---|---|
| Highest independent text-reasoning score | GLM-5.3 | It leads the current Intelligence Index comparison by 12 points. |
| Image, audio, or video understanding | Gemini 3.1 Pro Preview | Gemini accepts all three; base GLM-5.3 is text-only. |
| Fast, concise agent output | Gemini 3.1 Pro Preview | It generated faster and used far fewer output tokens in the independent evaluation. |
| Budget-sensitive prompts below 200K | Test total task cost | GLM has lower list prices, while Gemini had the lower measured task cost. |
| Prompts above 200K tokens | Model the full workload | Gemini switches to higher input, cache, and output rates. |
| Maximum generated output | GLM-5.3 | GLM documents 128K maximum output versus Gemini's 64K. |
| Production stability today | GLM-5.3 API | Gemini 3.1 Pro remains a preview endpoint; validate preview risk against your release policy. |
| Future self-hosting or model control | GLM-5.3 | Z.ai announced an open-weight release; verify availability and license before deployment. |
Frequently asked questions
Can I try GLM-5.3 before integrating it?
Use the OneInfer GLM-5.3 launcher to open a prepared prompt in the authenticated playground. Availability is checked against the current model catalog.
How should I treat benchmark claims?
Check the provenance label and harness version. Vendor-reported and independently verified results are deliberately shown as different evidence classes.
Is Gemini 3.1 Pro generally available?
No. Google documents the current API model as gemini-3.1-pro-preview and states that Gemini 3 models are currently in preview. Pin the model ID and regression-test behavior before production use.
Which model scored higher independently?
Artificial Analysis Intelligence Index v4.1.1 reports GLM-5.3 max at 60 and Gemini 3.1 Pro Preview at 48. That composite result should be supplemented with workload-specific multimodal and agent tests.
Which model is cheaper?
GLM-5.3 has lower standard first-party token prices. Gemini 3.1 Pro Preview nevertheless had a lower cost per task in the independent evaluation because it used substantially fewer output tokens. Gemini rates also increase above 200K prompt tokens.
Can both models process images, audio, and video?
No. Gemini 3.1 Pro Preview accepts text, image, audio, and video input. The base GLM-5.3 model accepts text only.
Put GLM-5.3 to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.