Skip to content

Science1 publisher3 min readPublished

GLM-5.3 says the quiet part: the base model did not change, the post-training did

Z.ai claims frontier agentic-coding scores at about 750B parameters, a third of Kimi K3, from extended post-training on the GLM-5.2 base. Open weights are promised in two weeks.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Z.ai announced GLM-5.3, currently only available in the coding plan, coming soon to their API and in two weeks' time to Hugging Face as open weights.
  • GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training.
  • The Z.ai blog post starts with the sentence: "Scaling post-training is all we did for GLM-5.3."
  • GLM-5.3 has approximately 750B parameters, described as a third of Moonshot AI's Kimi K3.
  • On many benchmarks GLM-5.3 has surpassed Moonshot AI's Kimi K3, and on some it has surpassed Claude Fable 5 or GPT-5.6-Sol.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

Z.ai has announced GLM-5.3, available first inside its coding plan, with API access to follow and open weights on Hugging Face in two weeks [1]. The interesting part is not the benchmark table but the provenance: Z.ai says GLM-5.3 uses the same base model as GLM-5.2, with substantially extended post-training, and its blog post opens with the line "Scaling post-training is all we did for GLM-5.3" [2][3].

The size claim is what should change planning assumptions. The model runs at roughly 750B parameters, which the newsletter Interconnects describes as a third of Moonshot AI's Kimi K3 [4], implying a Kimi K3 parameter count near 2.25T [1]. At two bytes per parameter, that is about 1.5 TB of weights to hold versus about 4.5 TB [2]. Interconnects reports GLM-5.3 surpassing Kimi K3 on many benchmarks and, on some, Claude Fable 5 or GPT-5.6-Sol, putting it more or less at the frontier of agentic coding evaluations [5][6]. Those are vendor-adjacent scores on an unreleased set of weights, not independent evals.

On method, Z.ai says it used more environments, more diverse tasks, and more compute spent training on them [7]. Interconnects reads this as an RL-dominated regime and argues distillation from Western frontier labs is not the major factor, on the grounds that you cannot distill RL environments, the infrastructure to run them at scale, or the algorithms to mix them [8][9]. The author does note a recent paper showing simple methods for extracting reasoning traces from frontier models, something Chinese labs could use at scale, and says it does not add up that U.S. labs have not patched the behavior faster while asking government for policy help [10][11]. Benchmaxxing, defined in the piece as tuning a model toward test sets so real-world performance diverges from paper scores, is raised and not the explanation offered [12].

The duller explanation is tenure. Zhipu AI was founded in 2019 [13]; GLM shipped in March 2021 from THUDM, Tsinghua University's data mining and knowledge engineering group [14]; GLM-130B in August 2022 [15]; ChatGLM on March 14, 2023, then ChatGLM2 on June 25 and ChatGLM3 on October 27 [16][17]; GLM-4 on January 16, 2024, with open-weight GLM-4-9B in June [18]; GLM-5 on February 11, 2026 [19]. That is nearly five years of continuous iteration on one model line before this release [3]. GLM-5.2, released June 22 of this year, was still in regular use by researchers weeks later for its speed and simplicity, with some deploying it on internal clusters to beat public serving latency [20].

Interconnects also makes a structural point operators should price in: Z.ai's time from finished model to release is likely days, while OpenAI and Anthropic take months of pre-release testing, so American labs probably hold better internal models while Chinese labs keep hillclimbing in public [21][22].

What to watch: whether the Hugging Face weights in two weeks are the same model as the coding-plan endpoint, whether independent agentic-coding harnesses reproduce the scores at 750B, and whether frontier labs close the reasoning-trace extraction path [1][4][10]. If the post-training-only story holds, the scarce input is environments and RL infrastructure, not pretraining compute.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories