Build4 publishers2 min readPublished Updated
Anthropic says a freely downloadable model builds exploits nearly as well as its restricted tool
Anthropic says Zhipu's freely downloadable GLM-5.3 built working V8 exploits in 50 of 410 tries, against 56 for its own restricted Claude Mythos Preview. With the weights public, its safeguards come off cheaply, so a lab that restricts its own model no longer keeps the capability out of reach.
The Engineer · Build desk

What happened
- GLM-5.3 refused openly malicious commands, but framed as a red-team exercise it attempted the target connection in 64 percent of runs, rising to 92 percent with prefilled reasoning and 100 percent once its refusals were removed.
- Paired with a human expert for a single day, the model found unknown flaws in a widely used browser's JavaScript engine and chained them into a web page that read any file on a visitor's machine, taking a private SSH key in the test.
- Vetted defenders given early access to Anthropic's Claude Mythos Preview through Project Glasswing have since found more than 10,000 vulnerabilities in critical software, the company says.
- The US agency CAISI independently judged GLM-5.3 the most cyber-capable open-weight model to date and put it about four months behind the best US systems.
- Using the smaller GLM-5.3-Flash, Anthropic combined a freshly disclosed Chrome bug with another known vulnerability into a reliable attack that bypassed a processor security feature, in a run that cost $20.40 in API time.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A capable exploit model now circulates as open weights, so a lab's choice to restrict its own release can no longer decide who obtains the capability.
- cost Safeguards added before release do not survive the download; Anthropic showed the refusal training comes off while science and cyber scores stay, so a publisher's safety work protects no one running the stripped weights.
- precedent Unlocked builds appeared within days of GLM-5.3's launch, so the next open-weight model at this capability should be expected to circulate with its safeguards stripped almost at once.
- contradiction The alarm comes from Anthropic, which sells the restricted alternative and, the-decoder notes, has its own reasons to raise it; the corroborating CAISI assessment tested US models with cyber safeguards off and counted systems only vetted users can reach, so both comparisons tilt toward restriction.
GLM-5.3 builds complete cyber exploits on its own, and unlike other models at that level it shipped without effective safeguards [3]. Those safeguards do not survive once the weights are public. Abliteration edits a model's parameters to remove the internal direction that produces a refusal, so it stops declining requests. Anthropic says this was its first time running the technique: about 2,200 GPU hours and roughly $4,400 in compute, with an estimate that a practiced team could repeat it for around $1,200 [21]. The evaluation never executed any code, so it cannot say whether a given attack would have landed, and the protected Claude models Anthropic tested stayed at zero throughout [20].
The two are close on capability. On ExploitBench, which targets known bugs in Chrome's V8 engine [8], GLM-5.3's success rate works out to about 12 percent against roughly 14 percent for Mythos Preview [1]. On its internal binary-exploitation set, the Frontier Red Team wrote that it evaluated several models on 100 randomly selected tasks and found that GLM-5.3 developed full control flow hijacks in 4% of trials, against 6% for Claude Mythos Preview. The team said that although GLM-5.3 performs below Mythos Preview there, a meaningful threshold has clearly been crossed, because earlier models such as Claude Opus 4.6 and GLM-5.2 succeeded on none of the tasks [26].
On release policy the two teams diverge. OpenAI is taking the same restricted approach with its Daybreak model, on the theory that vetted defenders getting the tool first is worth more than open access [7]. GLM-5.3 comes from Zhipu AI, which operates as Z.ai outside China [2]. Anthropic says it reported the browser vulnerabilities it found to the developers, and that findings in drivers and firmware are still under review [13].
Anthropic concludes that state and non-state actors will likely use models like GLM-5.3 to cause real harm, citing reports of attackers already using AI, and it argues that governments should test capable models [24].
What to watch
- Whether the browser vendor patches the flaws Anthropic reported, and what its driver and firmware findings turn out to be.
- Whether more unlocked GLM-5.3 variants spread beyond the ones released within days of launch.
- Whether any regulator moves to test or restrict open-weight models at this capability.