Product1 publisher3 min readPublished
Z.ai's GLM-5.3 beats Claude on CyberGym, then hands out the weights
Z.ai says its new model tops CyberGym and leads open-source models on Terminal Bench 3.0. The weights go to Hugging Face within two weeks, which is the part security teams should read twice.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Chinese AI developer Z.ai Co. debuted GLM-5.3, an open-source large language model, on August 14, 2026.
- GLM-5.3 is based on GLM-5.2, which Z.ai released in mid-July and which uses a mixture-of-experts architecture with 753 billion parameters and a 1 million token context window.
- GLM-5.3 has an identical design to GLM-5.2 but went through a more extensive post-training process.
- GLM-5.3 achieved the highest score of any open-source AI model on Terminal Bench 3.0, a benchmark that measures command line scripting capabilities.
- GLM-5.3 outperformed Anthropic's Claude Mythos 5 on CyberGym, a benchmark that evaluates the ability of LLMs to find code vulnerabilities.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
Z.ai released GLM-5.3 on Thursday, an open-source large language model built on the same architecture as July's GLM-5.2 but with a heavier post-training run [1][2][3]. The consequential detail is not the coding scores: it is that Z.ai says the model beat Anthropic's Claude Mythos 5 on CyberGym, a benchmark for finding code vulnerabilities, and that the weights are due on Hugging Face under an open-source license within two weeks [5][12].
The engineering is deliberately unglamorous. GLM-5.3 keeps GLM-5.2's mixture-of-experts design at 753 billion parameters with a 1 million token context window [2][3]. What changed is training: Z.ai says it put the model into sandboxes built to imitate developer workstations, with some exercises running for days, which it credits for improved long-horizon task performance [8]. The sandboxes themselves were generated by AI agents modelled on real software projects, with a separate judge agent verifying each challenge was solvable before the model saw it [9]. Reward signals were produced by automated pipelines [10], and the stack sits on two open-source components, slime for moving models from training to inference infrastructure and SAO for asynchronous reinforcement learning [11].
The claimed results, per Z.ai as reported by SiliconANGLE: the top open-source score on Terminal Bench 3.0, which tests command line scripting, and a 50 percent improvement over GLM-5.2 on an internal coding-agent benchmark [4][6]. Internal benchmarks are marketing until someone else runs them, but Terminal Bench and CyberGym are not Z.ai's to grade. The honest caveat is that the CyberGym win is one result out of three: GLM-5.3 trailed Anthropic's flagship on two other cybersecurity benchmarks, which the source does not name [7].
The field data is the more interesting number. Z.ai says the model has found more than 2,400 vulnerabilities across 269 software projects, with about half rated medium severity or higher [13][14]. That is roughly 1,200 medium-or-worse findings [15] and an average of about nine per project [16]. One flaw sits in code written 40 years ago, meaning code from around 1986 [17][18]. None of that has been independently verified, and the source gives no disclosure timeline or CVE references.
For anyone weighing whether to keep paying frontier API rates for agentic coding work, this is the strongest open-weights case so far, with two asterisks. First, GLM-5.3 is currently only reachable through Z.ai's GLM Coding Plan subscription; the weights are promised, not shipped [12]. Second, the announcement contains no pricing, so the cost argument is inference economics you have to model yourself [19].
The same release is a security planning problem. Once weights are published under an open license, the model runs on whatever hardware its operator owns, outside the subscription, the rate limits and the logging that come with a hosted API [20]. A vulnerability-discovery capability that Z.ai says is competitive with a frontier model becomes available to anyone who can rent GPUs, and the 2,400 findings are a demonstration of throughput, not a one-off [13].
What to watch: whether the weights actually land inside two weeks and under which license terms [12]; whether the two lost cybersecurity benchmarks are ever named [7]; and whether any of the 2,400 vulnerabilities surface as public advisories with the model credited [13].