Leadership1 publisher2 min readPublished
Anthropic finds attackers can bypass GLM-5.3's exploit safeguards up to 100% of the time
Anthropic says Zhipu AI's downloadable GLM-5.3 builds cyber exploits on its own, with safeguards that fail against simple attacks up to 100% of the time. It puts a capability once confined to gated frontier models within anyone's reach.
The Board Room · Leadership desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- NIST's Center for AI Standards and Innovation, in a Sept. 17 assessment, called GLM-5.3 the most cyber-capable open-weight model released to date and put it about four months behind the US frontier.
- On ExploitBench, which tests exploitation of known flaws in the V8 engine behind Google Chrome, GLM-5.3 wrote working end-to-end exploits in 50 of 410 attempts, close to Claude Mythos Preview's 56.
- On Anthropic's binary-exploitation benchmark GLM-5.3 landed a full control-flow hijack in 4 percent of 100 tasks against 6 percent for Claude Mythos Preview, while earlier Claude Opus 4.6 and GLM-5.2 scored zero.
- Anthropic had released its comparable model, Claude Mythos Preview, only through Project Glasswing, a limited channel that let trusted defenders find more than 10,000 vulnerabilities before similar models reached attackers.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- exposure Anyone who can download a model can now run autonomous exploit-building, a capability that until now sat with labs and the vetted users allowed onto guarded frontier systems.
- decision The head start Anthropic created by holding its own model back is closing, so defenders now have to plan patching against attackers who can run the same class of tool.
- capability Anthropic says the capability cuts both ways, giving defenders a tool to find and fix flaws in their own systems before attackers reach them.
The analysis comes from Anthropic, which builds the Claude models it compares GLM-5.3 against. [18] A company grading another lab's model is a reason to check the work, and an independent check exists: Anthropic says its findings broadly match a separate US government evaluation of GLM-5.3's capability. [8] The government evaluators are not selling a model, so their agreement is worth more than Anthropic's claim on its own.
Capability rankings and access are separate questions. In that government comparison, the US models were tested with their cyber safeguards switched off, and the frontier set included versions released only to vetted users. [9] So the ranking describes what well-guarded labs can produce under controlled conditions.
The clearest difference between the two models is in the safeguards. Anthropic says the same simple techniques that stripped GLM-5.3's protections failed against safeguarded Claude models in the same tests. [4]
Anthropic concentrated on exploit development because that is where Claude Mythos Preview showed a notable jump over earlier Claude models. [19] It says end-to-end exploit development is the capability that matters most to attackers. [20] It did not rely on automated scores alone. It put GLM-5.3 in front of human experts who did not know of existing vulnerabilities in the targets and asked them to find and exploit novel flaws with the model. The runs used sandboxed environments limited to offline targets, and they typically lasted a day or less. [17] Those runs measure what a person with the model could do in a short window.
What to watch
- Whether Zhipu AI adds misuse safeguards to GLM-5.3 or limits how the weights are distributed.
- Whether Anthropic or CAISI publish the full results of the human-expert exploit tests.
- Whether US frontier labs keep gating access as open-weight models close the capability gap.