GLM 5.3 Flash Is Out: A $0.15 Open-Weight Coding Model That Beats GPT-5.6 Luna
Z.ai just released GLM 5.3 Flash, an open-weight coding model scoring #3 on the Intelligence Index at $0.15/M input. What heavy AI users pay for cheap coding now.
Two days after Z.ai shipped the full GLM 5.3, it dropped a second weapon aimed straight at your monthly bill: GLM 5.3 Flash. Released on August 26, 2026, this open-weight model scores third on Artificial Analysis’s Intelligence Index at 57, ahead of GPT-5.6 Luna and DeepSeek V4 Flash, while its API pricing lands at a fraction of the US frontier. For anyone spending $300 or more a month on AI tokens, the number to memorize is simple: GLM 5.3 Flash costs $0.15 per million input tokens on the native Z.ai API, and even less on OpenRouter.

What GLM 5.3 Flash Actually Is
GLM 5.3 Flash is the fast, cheap sibling in the new GLM 5.3 line. It is a mixture-of-experts model with 320 billion total parameters and only 18 billion active, which is why it can serve tokens so cheaply while keeping a 1 million token context window and native reasoning. It is also genuinely multimodal on the input side, accepting text, image, and video, and outputting text, which makes it unusually versatile for agent pipelines that process screenshots, error traces, or video feeds as part of a coding loop.
The headline number is the Artificial Analysis Intelligence Index at 57, ranked third out of 109 models. That puts it ahead of GPT-5.6 Luna (52.3) and DeepSeek V4 Flash (51.8) on the same aggregate measure, in the same bracket as models that cost orders of magnitude more per token. On a few agentic and coding benchmarks, it holds its own against models priced ten to twenty times higher.
For heavy users the more important detail is that GLM 5.3 Flash is open-weight. That word carries two cost implications at once. First, it caps what Z.ai can charge you, because a competitor can run the same weights on cheaper GPU kits tomorrow and undercut the API. Second, it gives you a real self-hosting fallback if your API bill grows faster than the open-weight pricing resets the market.
The GLM 5.3 Flash Pricing That Changes Your Budget
Here is the exact GLM 5.3 Flash pricing as of late August, and how it stacks against the models it beats on intelligence:
- GLM 5.3 Flash (Z.ai native): $0.15 input, $0.50 output, $0.09 cached input, 83% cache discount, 1M context
- GLM 5.3 Flash (OpenRouter): $0.075 input, $0.25 output, $0.015 cache read, 1M context
- GPT-5.6 Luna: $0.20 input, $1.20 output
- DeepSeek V4 Flash: $0.22 input, $0.66 output (off-peak cache-miss)
- GPT-5.6 Sol: $4 input, $20 output
- Claude Opus 5: around $5 input, $25 output
Weave in the cache discount and the numbers get even better for the workloads heavy users actually run. Agentic coding is repetitive by nature, the same files, the same system prompt, the same tool definitions loaded hundreds of times a day. On OpenRouter, a warm cache read costs just $0.015 per million tokens. That is the cheapest cache tier in the frontier class right now, and it is precisely where your real-world effective rate lands once caching is on.

What the Cheap Frontier Class Means for Your Bill
The practical question for a heavy AI user is not whether GLM 5.3 Flash is competitive, it is where to slot it into a routing strategy. The economics are now genuinely lopsided in its favor for high-volume, lower-critically agentic coding.
Run the math on an output-heavy session that drains 500K output tokens. At GPT-5.6 Sol’s $20 per million that is $10 a session. At GLM 5.3 Flash’s $0.50 per million on the native API it is $0.25. Across a week of heavy agentic work that compounds into real money, the exact kind of line item TokenKarma tracks per provider.
The honest caveat is verbosity and raw capability. Artificial Analysis flags GLM 5.3 Flash as very verbose, generating around 150 million output tokens across its intelligence test suite against a median of 100 million. For an output-token-driven billing model, verbosity eats into your effective savings, and on the hardest agentic reasoning benches it still trails Sol and Claude Opus 5 by a visible margin. It also runs slower than average at about 50 tokens per second, which matters when latency, not just price, is the bottleneck.
That is why the right mental model is tiered routing, not replacement. Keep the frontier flagships for the hardest long-horizon reasoning, and route the bulk, high-volume, repetitive coding work to GLM 5.3 Flash. Because it is open-weight, the self-hosting option only strengthens that case over time.

The Playbook for Heavy AI Users
Concretely, this is how you should act on GLM 5.3 Flash over the next week:
1. Test it on your real workload, not a benchmark. The Intelligence Index is impressive, but your codebase is what matters. Run GLM 5.3 Flash on a slice of the agentic tasks you do daily and compare quality and token burn against Sol and Opus 5.
2. Quantify the verbosity tax. Because the model is very verbose, measure your actual effective output cost per task, not the list price. Set a token budget per agent run and watch whether GLM 5.3 Flash stays under it.
3. Turn caching on and keep your context stable. The $0.015 cache read on OpenRouter is the biggest lever. The more stable your system prompt and context window, the closer your blended rate gets to that floor. If you churn context constantly, you forfeit most of the discount.
4. Slot it into a tiered router. Route high-volume, lower-critically coding to GLM 5.3 Flash and reserve Sol, Opus 5, or Fable 5 for the hardest reasoning. A $10 and a $0.25 per-session model should not share a routing rule.
5. Watch the open-weight self-hosting ceiling. GLM 5.3 Flash is open-weight, which caps sustained price increases and gives you a real fallback. Budget a one-time self-hosting proof of concept now so you are not negotiating from a corner later.
The AI price war keeps rippling, and GLM 5.3 Flash is one of the clearest cost signals yet. An open-weight model that beats two proprietary frontier tiers on an aggregate intelligence index, at $0.15 on the native API and $0.075 on OpenRouter, resets what a heavy AI user should happily pay for bulk coding. Track it like every other provider, keep routing to whoever is cheapest for the task, and GLM 5.3 Flash just made that decision a lot easier.
Now available
Stop guessing your AI limits
The Mac app and web dashboard watch your Claude, ChatGPT, Gemini and more, and warn you before quotas hit.