Invest1 distinct publisher3 min readUpdated
Z.ai says its 743B-parameter GLM-5.3 hits 34.5% on its own code bench using 22% fewer output tokens than GLM-5.2. The weights are still two weeks out.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
Z.ai says its 743B-parameter GLM-5.3 hits 34.5% on its own code bench using 22% fewer output tokens than GLM-5.2. The weights are still two weeks out.
Z.ai, the Beijing lab formerly called Zhipu AI, shipped GLM-5.3 on Thursday through its GLM Coding Plan subscription and ZCode, pitching it as the "most capable open-weights model for coding" [1][2]. For anyone sizing an engineering org's model budget, the number that matters is not the 743 billion parameters but the token bill: Z.ai says the model reaches 34.5% on its in-house Z.ai Code Bench at Max effort while burning roughly 75,000 output tokens per task, against GLM-5.2's 23.4% at 96,000 [3][5].
That is about 22% fewer output tokens for a score 47% higher in relative terms [1][7]. Put differently, cost per point of benchmark performance falls from roughly 4,100 output tokens to roughly 2,200, a 47% reduction [2]. Z.ai's own framing is unusually modest about method: "Scaling post-training is all we did for GLM-5.3," the launch post says, describing a month of more environments, more tasks and more compute on the GLM-5.2 stack [4].
Two caveats before the procurement conclusion. The efficiency figures come from a benchmark Z.ai built, and the company concedes GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort, a five-point gap, while claiming a win over Claude Opus 4.8 on token economy [6][3]. And the open weights are not open yet: API access and downloadable weights follow staged safety evaluations, with the launch post putting weights about two weeks out [12].
On boards Z.ai does not own, the gap is real but narrow. GLM-5.3 scores 28.3 on Terminal Bench 3.0, 5.4 behind Fable 5 at 33.7 and 6.3 behind GPT-5.6 Sol at 34.6 [7][4]. On DeepSWE v1.1, which grades end-to-end fixes of real GitHub issues, GLM-5.3's 66.9 trails Fable 5's 69.7 by 2.8 points and open rival Kimi K3's 67.5 by 0.6 [8][5]. Decrypt's write-up nonetheless describes GLM-5.3 as beating Kimi K3 on the most relevant benchmarks, which is not what the DeepSWE line shows [9].
The security results are the loudest claim in the release and the likeliest reason weights ship late. Z.ai says GLM-5.3 leads CyberGym at 84.5% and more than doubles GLM-5.2 on exploitation benchmarks [10], and that the model flagged 2,436 vulnerabilities across 269 open-source projects, 1,097 of them medium-to-high severity [11].
Price is what turns a few benchmark points into a deferral decision. The GLM Coding Plan runs on a points quota with off-peak calls at half rate, and Zhipu's API sits at roughly a tenth of US frontier per-token pricing; GLM-5.2's official rate was $1.40 in and $4.40 out per million tokens [13]. GPT-5.3-Codex lists at $1.75 and $14, so 1.25x the input and 3.2x the output of the older GLM, with Claude Opus 4.8 near the top of Anthropic's tiers [14][6]. A team six points off the leaderboard at a third to a tenth of the token cost, using fewer tokens per task, has a defensible reason not to sign a multi-year closed-model commitment this quarter. Chinese open-weight models already out-consume American ones on OpenRouter token usage [16].
Watch three things: whether the weights actually land in two weeks and in full, whether independent runs reproduce the 75,000-token figure outside Z.ai's harness, and whether the CyberGym and vulnerability-discovery numbers change what a US Entity List lab is willing to publish at all [12][5][10][15].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Z.ai calls GLM-5.3 the "most capable open-weights model for coding".
Z.ai says GLM-5.3 clears 34.5% on its in-house Z.ai Code Bench at Max effort while using roughly 75,000 output tokens per task, against GLM-5.2's 23.4% at 96,000 output tokens.
Z.ai's GLM Coding Plan runs on a points quota with off-peak calls costing half, and Zhipu's API is priced at roughly a tenth of US frontier per-token rates; GLM-5.2's official rate was $1.40 in and $4.40 out per million tokens.
GPT-5.3-Codex is priced at $1.75 in and $14 out per million tokens, and Claude Opus 4.8 sits near the top of Anthropic's tiers.
Chinese AI lab Z.ai (formerly Zhipu AI) released GLM-5.3 on Thursday through its GLM Coding Plan subscription and ZCode.
GLM-5.3 is a 743-billion-parameter model built by scaling post-training on the GLM-5.2 base.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single outlet relaying vendor-reported numbers
Every capability, efficiency, security and pricing figure traces to Z.ai's launch post and X post as summarized by one publisher. The headline efficiency result is measured on the vendor's own in-house Code Bench, there is no independent replication of Terminal Bench, DeepSWE or CyberGym scores, and the weights that would let anyone verify are not released. Release, channel and Entity List facts are solid; the performance record is attributed rather than corroborated.
Shipped to one paid channel, weights pending
There is a dated, concrete release event, but distribution is confined to the GLM Coding Plan subscription and ZCode, with API and downloadable weights staged behind safety evaluations about two weeks out. The supplied material contains no user counts, deployment cases, or verified usage share — the OpenRouter framing is an unquantified assertion — so measurable uptake for GLM-5.3 itself is minimal.
Superlative framing outruns the tables
The launch is billed as the "most capable open-weights model for coding" and the report adds that GLM-5.3 beats Kimi K3 on the most relevant benchmarks, yet the article's own numbers put Kimi K3 ahead on DeepSWE v1.1 and closed models ahead on Terminal Bench 3.0 and Z.ai's Code Bench. The open-weights claim describes weights that will not exist publicly for about two weeks. The genuine, well-specified result — roughly 22% fewer output tokens per task and about half the tokens per score point — is real but narrower than the framing.
Vendor launch material driving a paid plan
The story's factual base is a promotional launch post and X post from a lab that monetizes the model through a subscription with a points quota, and that benefits commercially from a price-versus-U.S.-frontier comparison and from an open-weights reputation ahead of any actual weight drop. The benchmark that carries the headline result is the vendor's own. The publisher reproduces this framing with minimal adversarial sourcing, and the Entity List context gives Z.ai an additional reason to emphasize independence from U.S. suppliers.
Low — one publisher, one vendor voice
Release mechanics, access channels, pricing framing and Entity List status are reliably reported and internally consistent, which supports moderate confidence in the shape of the story. Confidence in the performance and security claims is low: single publisher, single vendor origin, in-house benchmark, no weights to test, an internal contradiction over Kimi K3, and an unsupported usage assertion.
build
GLM-5.3 kept the base model and bought ten times the environments instead2 distinct publishers
product
Z.ai's GLM-5.3 beats Claude on CyberGym, then hands out the weights1 distinct publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
build
GLM-5.3 is a paper, not an endpoint: Z.ai publishes research before weights1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026