Published · 6d agoSecurity2 min read
GLM-5.3's 14-day fuse: a 2,436-bug discovery model heads for open weights
Zhipu released GLM-5.3 through its coding service on August 14 and says weights follow in about two weeks. The vulnerability count behind it is Zhipu's own figure.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- Zhipu AI released GLM-5.3 on August 14, 2026 through its GLM Coding Plan and claimed the model is the strongest open-weights coding system currently available.
- GLM-5.3 model weights are not yet downloadable; Zhipu said it plans to release them in approximately two weeks after security reviews, so independent benchmark and tester results are not yet available.
- Zhipu's 50% improvement claim and its 'strongest open-weights coding model' description rest on the company's own evaluations; the reviewed release materials do not provide an independently reproduced benchmark set for GLM-5.3.
- Zhipu said GLM-5.3 scored 84.5 per cent on CyberGym, a benchmark measuring whether models can identify and validate security flaws from source code, above Anthropic's Mythos 5 at 83.8 per cent and OpenAI's GPT-5.6 Sol at 83.6 per cent.
- On ExploitBench, which gauges how far AI models climb the exploitation ladder, GLM-5.3 scored 54.4 per cent, trailing Mythos 5 at 78 per cent and GPT-5.6 Sol at 76.5 per cent.
Compiled by The WatchSomething wrong?How this is made
Why it matters
Zhipu put GLM-5.3 behind its GLM Coding Plan on August 14, 2026, and said the weights follow in roughly two weeks, once security reviews finish [1][2]. That interval is the story: for about fourteen days, a model the company credits with 2,436 vulnerabilities across 269 projects exists only as a service somebody can throttle [8][20].
What the figure turns on is which half of the offensive workflow travelled. On CyberGym, which tests whether a model can identify and validate flaws from source code, Zhipu reports 84.5 per cent, ahead of Anthropic's Mythos 5 at 83.8 and OpenAI's GPT-5.6 Sol at 83.6 [4]. On ExploitBench, which measures how far a model climbs the exploitation ladder, it reports 54.4 per cent against 78 and 76.5, leaving it 23.6 points behind Mythos [5][15]. So the capability about to become un-gatable is the finding, not the finishing. The trajectory is the part worth noting: the ExploitBench score more than doubled from GLM-5.2's 24.4 per cent, and on ExploitGym the model went from 29 completed tasks in two hours to 105, on the same base model with scaled post-training and added vulnerability-discovery data [6][7][12].
The ledger is where the two weeks bite. Zhipu lists 107 critical and 990 high-severity findings, which is exactly its 1,097 medium-to-high total; 53 are published and 2,383 remain under embargo, spanning kernels, operating systems, browser engines, open-source infrastructure and network protocols [9][18][17]. The company has not said how many were previously unknown, or how many anyone independently reproduced [11]. All of the scores are Zhipu's own evaluations, with no independently reproduced benchmark set available yet [3]. Neil Shah of Counterpoint Research put the constraint plainly: once weights are released freely, built-in guardrails can be stripped without any cognizance or control [13].
For scale on the defender side, CISA's implementation guidance for BOD 26-04 allots agencies the first two hours after a KEV addition to decide whether one CVE meets the under-three-days remediation threshold [16].
Watch the license and model card when the files publish, and whether the 2,383 embargoed entries begin moving through disclosure once anyone can rerun the discovery pipeline locally [9][2].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Zhipu AI released GLM-5.3 on August 14, 2026 through its GLM Coding Plan and claimed the model is the strongest open-weights coding system currently available.
- [2]
GLM-5.3 model weights are not yet downloadable; Zhipu said it plans to release them in approximately two weeks after security reviews, so independent benchmark and tester results are not yet available.
ReportedView cited source - [3]
Zhipu's 50% improvement claim and its 'strongest open-weights coding model' description rest on the company's own evaluations; the reviewed release materials do not provide an independently reproduced benchmark set for GLM-5.3.
ReportedView cited source - [4]
Zhipu said GLM-5.3 scored 84.5 per cent on CyberGym, a benchmark measuring whether models can identify and validate security flaws from source code, above Anthropic's Mythos 5 at 83.8 per cent and OpenAI's GPT-5.6 Sol at 83.6 per cent.
- [5]
On ExploitBench, which gauges how far AI models climb the exploitation ladder, GLM-5.3 scored 54.4 per cent, trailing Mythos 5 at 78 per cent and GPT-5.6 Sol at 76.5 per cent.
ReportedView cited source - [6]
GLM-5.3's ExploitBench score more than doubled from GLM-5.2's 24.4 per cent, according to Zhipu.
ReportedView cited source
Sources & coverage · 4 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.



