Invest1 publisher3 min readPublished Updated
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead
Z.ai says its 743B-parameter GLM-5.3 hits 34.5% on its own code bench using 22% fewer output tokens than GLM-5.2. The weights are still two weeks out.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Chinese AI lab Z.ai (formerly Zhipu AI) released GLM-5.3 on Thursday through its GLM Coding Plan subscription and ZCode.
- Z.ai calls GLM-5.3 the "most capable open-weights model for coding".
- GLM-5.3 is a 743-billion-parameter model built by scaling post-training on the GLM-5.2 base.
- Z.ai wrote: "Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack... Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them."
- Z.ai says GLM-5.3 clears 34.5% on its in-house Z.ai Code Bench at Max effort while using roughly 75,000 output tokens per task, against GLM-5.2's 23.4% at 96,000 output tokens.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
Z.ai, the Beijing lab formerly called Zhipu AI, shipped GLM-5.3 on Thursday through its GLM Coding Plan subscription and ZCode, pitching it as the "most capable open-weights model for coding" [1][2]. For anyone sizing an engineering org's model budget, the number that matters is not the 743 billion parameters but the token bill: Z.ai says the model reaches 34.5% on its in-house Z.ai Code Bench at Max effort while burning roughly 75,000 output tokens per task, against GLM-5.2's 23.4% at 96,000 [3][5].
That is about 22% fewer output tokens for a score 47% higher in relative terms [1][7]. Put differently, cost per point of benchmark performance falls from roughly 4,100 output tokens to roughly 2,200, a 47% reduction [2]. Z.ai's own framing is unusually modest about method: "Scaling post-training is all we did for GLM-5.3," the launch post says, describing a month of more environments, more tasks and more compute on the GLM-5.2 stack [4].
Two caveats before the procurement conclusion. The efficiency figures come from a benchmark Z.ai built, and the company concedes GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort, a five-point gap, while claiming a win over Claude Opus 4.8 on token economy [6][3]. And the open weights are not open yet: API access and downloadable weights follow staged safety evaluations, with the launch post putting weights about two weeks out [12].
On boards Z.ai does not own, the gap is real but narrow. GLM-5.3 scores 28.3 on Terminal Bench 3.0, 5.4 behind Fable 5 at 33.7 and 6.3 behind GPT-5.6 Sol at 34.6 [7][4]. On DeepSWE v1.1, which grades end-to-end fixes of real GitHub issues, GLM-5.3's 66.9 trails Fable 5's 69.7 by 2.8 points and open rival Kimi K3's 67.5 by 0.6 [8][5]. Decrypt's write-up nonetheless describes GLM-5.3 as beating Kimi K3 on the most relevant benchmarks, which is not what the DeepSWE line shows [9].
The security results are the loudest claim in the release and the likeliest reason weights ship late. Z.ai says GLM-5.3 leads CyberGym at 84.5% and more than doubles GLM-5.2 on exploitation benchmarks [10], and that the model flagged 2,436 vulnerabilities across 269 open-source projects, 1,097 of them medium-to-high severity [11].
Price is what turns a few benchmark points into a deferral decision. The GLM Coding Plan runs on a points quota with off-peak calls at half rate, and Zhipu's API sits at roughly a tenth of US frontier per-token pricing; GLM-5.2's official rate was $1.40 in and $4.40 out per million tokens [13]. GPT-5.3-Codex lists at $1.75 and $14, so 1.25x the input and 3.2x the output of the older GLM, with Claude Opus 4.8 near the top of Anthropic's tiers [14][6]. A team six points off the leaderboard at a third to a tenth of the token cost, using fewer tokens per task, has a defensible reason not to sign a multi-year closed-model commitment this quarter. Chinese open-weight models already out-consume American ones on OpenRouter token usage [16].
Watch three things: whether the weights actually land in two weeks and in full, whether independent runs reproduce the 75,000-token figure outside Z.ai's harness, and whether the CyberGym and vulnerability-discovery numbers change what a US Entity List lab is willing to publish at all [12][5][10][15].