Published · 6d agoBuild2 min read
84.5% on CyberGym is a post-training number, and Z.ai's docs print no score
Z.ai says GLM-5.3 shares GLM-5.2's base model and credits every gain to post-training. That makes the cyber result a claim about training environments, not a bigger model.
Written for builders.See today for builders

What happened
- GLM-5.3 uses the same base model as GLM-5.2, with all improvements driven by post-training, according to Z.ai's developer documentation.
- Z.ai's GLM-5.3 overview states the model 'achieves the best performance to date on the CyberGym vulnerability discovery benchmark' and describes cybersecurity gains as improving faster than expected as post-training scale expands; it publishes no numeric CyberGym score on the page.
- Z.ai states that GLM-5.3's scores on vulnerability exploitation benchmarks exceed twice those of GLM-5.2, and that the deeper the model progresses along the vulnerability exploitation chain, the more pronounced its gains over GLM-5.2 become.
- Z.ai reports GLM-5.3 achieving a 50% performance gain over GLM-5.2 on Z.ai Code Bench, and state-of-the-art performance among open-source models on Terminal Bench 3.0 and Agents' Last Exam (CLI).
- GLM-5.3 supports text-only inputs with a 1M-token context window and 128K maximum output, always operates with reasoning enabled at low, high or max effort, and its model API is listed as 'available soon' rather than released.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
The 84.5% CyberGym figure attached to GLM-5.3 is vendor-reported, and worth being precise about: Z.ai's own overview page asserts only that the model "achieves the best performance to date" on the CyberGym vulnerability discovery benchmark, without publishing a numeric score [2]. What the documentation does commit to is more load-bearing than the digits. GLM-5.3 runs on the same base model as GLM-5.2, with all improvements attributed to post-training [1].
That is the mechanism the number turns on. If the base is unchanged, a jump in vulnerability discovery is not a scaling result; it is a claim about training environments and reward design. Z.ai says the same post-training push produced a 50% gain over GLM-5.2 on its internal Z.ai Code Bench and state-of-the-art open-source results on Terminal Bench 3.0 and Agents' Last Exam (CLI) [4], and that scores on vulnerability exploitation benchmarks exceed twice GLM-5.2's, with the margin widening the deeper the model goes along the exploitation chain [3]. The company describes environments built around production workflows, some representing several days of work for an experienced engineer, including ML infrastructure tasks with access to compute clusters, storage, internal documentation and prior experiment results [6].
Why that is hard is documented outside the vendor. The Agent-RLVR authors report that verifiable-reward RL loses efficacy in agentic settings because failure rates are high and the reward landscape is too sparse [9]. With 817 curated software engineering environments plus teacher-style guidance, they lifted Qwen-2.5-72B-Instruct from 9.4% to 22.4% pass@1 on SWE-Bench Verified, and to 27.8% with a test-time reward model [10] - roughly a 2.4x move on an untouched base [11]. Same shape of claim, independently reported.
What would move the CyberGym figure: a published score with harness and subset detail, a third-party rerun, and the API itself, which Z.ai lists as not yet released [5]. Until then it is one company's benchmark on one company's environments.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
GLM-5.3 uses the same base model as GLM-5.2, with all improvements driven by post-training, according to Z.ai's developer documentation.
- [2]
Z.ai's GLM-5.3 overview states the model 'achieves the best performance to date on the CyberGym vulnerability discovery benchmark' and describes cybersecurity gains as improving faster than expected as post-training scale expands; it publishes no numeric CyberGym score on the page.
- [3]
Z.ai states that GLM-5.3's scores on vulnerability exploitation benchmarks exceed twice those of GLM-5.2, and that the deeper the model progresses along the vulnerability exploitation chain, the more pronounced its gains over GLM-5.2 become.
- [4]
Z.ai reports GLM-5.3 achieving a 50% performance gain over GLM-5.2 on Z.ai Code Bench, and state-of-the-art performance among open-source models on Terminal Bench 3.0 and Agents' Last Exam (CLI).
- [5]
GLM-5.3 supports text-only inputs with a 1M-token context window and 128K maximum output, always operates with reasoning enabled at low, high or max effort, and its model API is listed as 'available soon' rather than released.
ReportedView cited source - [6]
Z.ai says GLM-5.3's post-training environments were pushed toward real units of expert work covering production workflows, some representing several days of work for an experienced engineer, for example an ML infrastructure task giving the model access to compute clusters, storage systems, internal documentation, codebases and experiment results.
Sources & coverage · 3 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- docs.z.ai6d agoGLM-5.3 - Overview - Z.AI DEVELOPER DOCUMENT

