tokenkarma is in beta. Your feedback shapes what ships next.
7 min read B2C power user

GLM 5.3 Beats GPT-5.6 Sol at Coding for 1/5 the Cost: What Heavy Users Pay Now

GLM 5.3 from Z.ai posts Terminal-Bench 3.0 at 28.3 and edges GPT-5.6 Sol, at $1.40/M input and $4.40/M output. What heavy AI users should pay for agentic coding now.

GLM 5.3 Beats GPT-5.6 Sol at Coding for 1/5 the Cost: What Heavy Users Pay Now

A new open-weights model landed in mid-August that is rewriting the cost math for agentic coding. GLM 5.3 from Z.ai (Zhipu AI) scores in the same Intelligence bracket as the flagship proprietary models, yet its API pricing lands at roughly a fifth of the OpenAI and Anthropic flagships. For anyone spending $300 or more a month on AI tokens, the number you need to know is simple: GLM 5.3 matches or beats GPT-5.6 Sol on several coding benchmarks while charging $1.40 per million input tokens and $4.40 per million output tokens, against $4 and $20 for Sol.

What Is GLM 5.3 and Why It Matters for Your Bill

GLM 5.3 launched on August 14, 2026, and it is built on the same base model as GLM 5.2. Every improvement in this release comes from post-training, not a new architecture, which is a meaningful detail for budget planning: the jump is real, but it is a software improvement on an existing model line, not a new silicon class.

What got the community’s attention on Hacker News is the combination of frontier-level coding scores and a small price tag. Z.ai reported Terminal-Bench 3.0 at 28.3, up from 4.6 on the previous release, and CyberGym at 84.5%, where it leads the open-weight field and edges past Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol. On the Ed-o-meter real-world suite, an independent 28-task harness, GLM 5.3 delivered a 100% pass rate at about a fifth of GPT-5.5’s cost for the same lap.

The Artificial Analysis Intelligence Index puts GLM 5.3 at 60, ranked 9th out of 187 models. That places it alongside the proprietary frontier while its cache discount of 81% (cached input at $0.68) and 1 million token context window make it unusually attractive for long, repetitive agentic workloads.

The Pricing Table Heavy Users Actually Care About

When you compare list prices for output-heavy coding, the gap is stark. These are the headline per-million-token rates as of late August 2026:

  • GLM 5.3 (max): $1.40 input, $4.40 output, $0.68 cached input, 81% cache discount, 1M context
  • GPT-5.6 Sol: $4 input, $20 output, $0.40 cached input (promotional through Nov 21, 2026)
  • Claude Opus 5 / Fable 5 Max: roughly $5 input, $25 output
  • Grok 4.6: $2 input, $6 output

Do the math on a session that drains 500K output tokens. At Sol’s $20 per million that is $10 per session. At GLM 5.3’s $4.40 per million it is $2.20. Across a month of heavy agentic work, that difference compounds into hundreds of dollars, the exact kind of line item TokenKarma users track weekly.

There are three caveats to weigh before you route your production workload to GLM 5.3.

Caveat One: The Scores Are Vendor-Reported (For Now)

Every headline benchmark above is a Z.ai claim. No independent lab has re-run the full set under a single harness. The weights were not publicly downloadable at launch; Z.ai said it would release them roughly two weeks after launch, in late August, pending safety evaluation. That means no one outside Z.ai could reproduce the numbers at first.

Independent signals that do not depend on Z.ai look good. The Ed-o-meter is a third-party real-world harness and it confirms the cheap-and-capable story. But a careful operator treats the vendor numbers as a best case until the weights land and the community re-tests.

Caveat Two: Verbosity Inflates Your Output Bill

GLM 5.3 is unusually verbose. During the Intelligence Index evaluation it generated 170 million output tokens against a median of 72 million. At $4.40 per million output, verbosity is a real cost driver, not a cosmetic quirk. A model that writes more tokens than its peers can erase much of its list-price advantage on output-heavy tasks.

If you route to GLM 5.3, set explicit reasoning effort bounds, cap output length at the API level, and monitor tokens per task rather than just cost per token. The smart play is to treat GLM 5.3 as a coding specialist for well-scoped tasks where verbosity is controllable, not as a blanket replacement for your whole stack.

Caveat Three: It Does Not Win Every Row

GLM 5.3 trails Fable 5 and GPT-5.6 Sol on offensive-security benchmarks and on several coding tests. “Beats the frontier” oversells a mixed picture. It leads on agentic coding and cyber-defense use cases but is not the top model for every workload. That matters for routing discipline: a young open-weight model still carries integration and reliability risk, the same burn-in concern the community flagged for Grok 4.6 in a fresh acquisition.

What This Means for Heavy AI Users

The strategic takeaway is about leverage, not just price. GLM 5.3 is the strongest signal yet that frontier-quality coding no longer requires flagship spending. For a heavy user the practical playbook is:

  1. Route token-heavy, repetitive coding to GLM 5.3 to capture the 4 to 5x list-price gap on output tokens.
  2. Keep a flagship fallback for offensive-security and the specialist rows where GLM 5.3 trails, and for anything where vendor-reported scores are not enough to bet on.
  3. Watch the weights release in late August to lock in option value and free yourself from a single vendor’s API, exactly as the open-weight hedge in the current price war rewards.
  4. Track tokens per task, not just price per token, to neutralize the verbosity penalty before it eats the savings.

The price war has now reached the open-weight tier with real teeth. GLM 5.3 does not replace the frontier, but it does what the best budget models always do: it forces every heavy user to re-run their routing math. If you are paying flagship prices for coding output, the $4.40 output rate is worth a serious test run before the weights ship and demand for the API pushes availability around.