Model facet · Pricing

GLM-5.3-Flash API pricing

GLM-5.3-Flash lists at $0.075 input and $0.250 output per 1M tokens as of 27 Aug 2026 on the OpenRouter Z.ai listing. It is 2.5× cheaper than DeepSeek V4 Flash 0731's *replacement* on input — but DeepSeek V4 Flash 0731 is still 2.5× cheaper than GLM-5.3-Flash on input and 3.3× on output. Qwen3.8-Flash is the multimodal flash option at 2.1× the input price.

Published API pricing per 1M tokens

Source: each model's own OpenRouter listing, fetched 27 August 2026. Sorted cheapest first on blended cost. Not reordered to flatter any model in the table.

ModelInput $/1MOutput $/1MBlended $/1MInput × GLMContext (tokens)MultimodalProviders on OpenRouterReleased
DeepSeek V4 Flash 0731$0.030$0.075$0.0410.4×1,310,720Not stated231 Jul 2026
GLM-5.3-Flash$0.075$0.250$0.1191.0×1,048,576Native126 Aug 2026
Qwen3.8-Flash$0.160$0.470$0.2382.1×1,000,000Image + text126 Aug 2026
Qwen3.8-Max$2.000$6.000$3.00026.7×1,000,000Yes13 Aug 2026

How blended cost is calculated

Blended cost assumes a 1:3 input/output token ratio (edit cell B3 to test other workloads). At 1:1 GLM-5.3-Flash blends to $0.163 and Qwen3.8-Flash to $0.315, widening the gap. At 1:9 it narrows. Treat blended as a workload sensitivity test, not a neutral sort key.

Ready to test the workflow?

Create account & add credits

Anchors built on "cheapest" or "budget" claims

Any anchor built on "cheapest", "budget" or "low cost" pointed at GLM-5.3-Flash is a claim the table above disproves. DeepSeek V4 Flash 0731 is 2.5× cheaper on input and 3.3× on output; Qwen3.8-Flash is 2.1× more expensive on input and 1.9× on output. The defensive position for GLM-5.3-Flash is native multimodal (image + text input) at flash-tier pricing — not the cheapest line item.

Frequently asked questions

How much does GLM-5.3-Flash cost per token?

$0.075 per 1M input tokens and $0.250 per 1M output tokens as of 27 August 2026. Blended at a 1:3 input/output ratio, $0.119 per 1M tokens.

Is GLM-5.3-Flash the cheapest flash-tier model?

No. DeepSeek V4 Flash 0731 is 2.5× cheaper on input and 3.3× on output. GLM-5.3-Flash's position is native multimodal flash-tier (image + text input) at $0.075 input — not the cheapest.

Where is the live GLM-5.3-Flash model page?

The canonical model page with current OneInfer pricing, capabilities, and availability is /models/zai-org/GLM-5.3-Flash. This page is a focused facet of that entity, not a replacement for it.

How should I treat benchmark or price claims?

Check each claim's provenance label and observed date. Vendor-reported and independently verified numbers are shown as separate evidence classes on this hub.

Put GLM-5.3-Flash to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.