Build1 publisher3 min readPublished
Open weights caught up on finding bugs. They did not catch up on using them.
Z.ai says GLM-5.3 edges Anthropic's restricted Mythos 5 at vulnerability discovery while losing badly at exploitation. On vendor numbers, the defensive half is commoditising first.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Mythos 5 is described as a version of Anthropic's Claude Fable 5 model with certain cybersecurity safeguards removed, and access is restricted to vetted organizations rather than released publicly.
- GLM-5.3 is a new general-purpose coding model from Chinese AI company Z.ai, released as open source, and Z.ai says it was not created purely as a cybersecurity product.
- Z.ai reported CyberGym scores of 84.5 percent for GLM-5.3 and 83.8 percent for Mythos 5.
- Z.ai reported ExploitBench scores of 54.4 percent for GLM-5.3 and 78.0 percent for Mythos 5.
- All comparative figures in the report are Z.ai's own reported results.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Z.ai says its open-source GLM-5.3 scored 84.5 percent on CyberGym against 83.8 percent for Anthropic's Mythos 5, but only 54.4 percent on ExploitBench against Mythos 5's 78.0 percent [3][4]. If those self-reported figures survive contact with independent testers, the cheap, downloadable half of AI security tooling is the half that finds problems, and the guarded half is still the one that turns a finding into something that works [5][9].
The asymmetry is the story. On discovery, the reported margin is 0.7 points in the open model's favour, which is noise [6]. On exploitation, GLM-5.3 lands at roughly 70 percent of Mythos 5's score, a 23.6 point deficit [7][8].
The time-budgeted attack numbers say something sharper. Under a two-hour budget, Z.ai reports 105 attack tasks for GLM-5.3 against 181 for Mythos 5; at six hours, 130 against 247 [10][11]. Tripling the clock adds 36.5 percent to the restricted model's total and 23.8 percent to the open one, so the open model's share of Mythos 5's output falls from 58.0 percent to 52.6 percent [12][13][14]. Longer runs do not close this gap; they widen it. That is the signature of a deficit in sustained multi-step execution rather than in code comprehension, which matches the pipeline the source lays out: read code, understand the system, find the bug, verify it, build the exploit, run the attack [15]. Recognising that a string-concatenated SQL query looks injectable is a different job from proving it is reachable given the application's authentication, input validation, and runtime behaviour [16][17].
For defenders, that split is usable. A model that can only reliably do discovery is still a model you can point at a repository on every commit instead of every quarterly audit [18][19]. Z.ai's other claim matters more for planning than the benchmark table: GLM-5.3 reportedly started from the same base model as GLM-5.2 and acquired its security capability through extended post-training, reinforcement learning, longer task environments, and more diverse security tasks, not a separate security-specific architecture [2][20]. If that is right, discovery capability is a post-training add-on to competent general coding models, which means it will keep arriving in open weights whether or not anyone plans a security release [21].
The caveats are load-bearing. Every number here is Z.ai's own reported result, relayed through a single write-up, with no independent reproduction cited [5][22]. Mythos 5 is described as a version of Anthropic's Claude Fable 5 with certain cybersecurity safeguards removed, distributed only to vetted organisations, which is exactly the access regime that makes third-party replication of the comparison hard [1][23]. A vendor benchmarking itself against a model most testers cannot obtain is a claim, not a measurement.
Watch three things. Whether anyone outside Z.ai reproduces the CyberGym and ExploitBench figures, and on which task subsets. Whether the two-hour-to-six-hour scaling penalty shrinks in the next open checkpoint, since that, not the discovery score, is the number that tracks end-to-end attack capability. And whether the vetted-access model for exploitation-capable systems holds once the discovery half is freely downloadable, because the restriction only buys time on the stage where the open models are still behind [9][23].