Anthropic says Zhipu's freely downloadable GLM-5.3 built working V8 exploits in 50 of 410 tries, against 56 for its own restricted Claude Mythos Preview. With the weights public, its safeguards come off cheaply, so a lab that restricts its own model no longer keeps the capability out of reach.
Perspective Coverage
4 publishers
- Builder
- Builder 41%
- Operator
- Operator 38%
- Investor
- Investor 21%
Reality
- Evidence62
- Adoption30
- Hype gap+15
- Incentives72
- Confidence62
VulnCheck counts Chrome CVEs up 563% and GitHub-issued CVEs up 476% this year, a rise it calls consistent with AI-assisted bug finding. Whether the volume lasts is unknown, so exploitation data still sets patch order.
Reality
- Evidence55
- Adoption45
- Hype gap+20
- Incentives55
- Confidence50
Anthropic says Zhipu AI's downloadable GLM-5.3 builds cyber exploits on its own, with safeguards that fail against simple attacks up to 100% of the time. It puts a capability once confined to gated frontier models within anyone's reach.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+20
- Incentives75
- Confidence50
Australia's Signals Directorate says attackers are using stolen AI API keys, tokens and hijacked sessions to get into organisations' AI services. Its guidance tells customers to protect those credentials themselves. In one reported case, a stolen key ran up about US$600,000 in model credits over three weeks.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence50
Anthropic's IPO prospectus gives a single Class F share, voted by its seven cofounders, 50.1% of votes on key corporate matters. Teams on Claude get a vendor public shareholders cannot outvote on those matters while the founder group holds together.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives65
- Confidence50
GitHub's volume counters and Chainguard's account of agent-written code both point at dependency review as the step nobody is doing. Mandiant estimates mean time-to-exploit at minus seven days in 2025.
Reality
- Evidence43
- Adoption55
- Hype gap+27
- Incentives76
- Confidence42
Anthropic says Claude Mythos Preview found and exploited previously unknown flaws in every major operating system and web browser during a month of testing. Its disclosure process keeps the rest unnamed until patches ship.
Publishers:red.anthropic.com
Reality
- Evidence36
- Adoption20
- Hype gap+35
- Incentives78
- Confidence55
Oracle put ChatGPT Enterprise and OpenAI's Codex in front of about 160,000 employees in April, and 80 percent were using them within three months. Then the bills arrived and, the CIO said, surprised the company.
Reality
- Evidence40
- Adoption68
- Hype gap+15
- Incentives68
- Confidence46
Luke Heath's written answers to Lets Data Science put eight named controls around the model, among them scoped credentials, engineer approval and hand reproduction of every candidate finding. Heath has not disclosed results yet.
Reality
- Evidence45
- Adoption20
- Hype gap−10
- Incentives70
- Confidence50
Project Glasswing surfaced an estimated 6,202 high or critical vulnerabilities in foundational open source, with 97 confirmed fixed in two months. What the confirmation count actually measures decides how much of that gap is real.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence28
Anthropic's red team says the scarce reverse-engineering skill that used to give defenders weeks is no longer the bottleneck. It measured that on Firefox and Windows kernel fixes, where a 19-day median gap counts as fast.
Reality
- Evidence55
- Adoption20
- Hype gap+22
- Incentives72
- Confidence48
A July weekend inside OpenAI's test environment produced the case study for runtime authority limits, and CrowdStrike is now selling both the layer that caps customer agents and the agents it wants trusted with production systems.
Reality
- Evidence36
- Adoption22
- Hype gap+34
- Incentives84
- Confidence42
Project Glasswing gives twelve launch partners and more than forty critical-infrastructure maintainers access to Claude Mythos Preview, which Anthropic says has already found thousands of high-severity vulnerabilities.
Reality
- Evidence30
- Adoption25
- Hype gap+45
- Incentives80
- Confidence60
Every number in the account traces to a single vendor blog post that lists no CVE identifiers and no disclosure dates. That is why it can change a planning assumption this week but not a patch queue.
Reality
- Evidence24
- Adoption18
- Hype gap+62
- Incentives82
- Confidence60
Anthropic sent 8% of its model's 23,019 findings out for review and stopped there, citing a shortage of people to check them. Where independent scoring did happen, 14 of 27 severity ratings had to be moved.
Reality
- Evidence58
- Adoption33
- Hype gap+42
- Incentives68
- Confidence55
Insikt Group counted 215 actively exploited CVEs in six months, up 34% year on year, and describes attackers running them through remote access utilities, package registries and payment flows defenders already permit.
Reality
- Evidence61
- Adoption63
- Hype gap−7
- Incentives71
- Confidence57
The benchmark hands an agent an input that already crashes a program and asks it to escalate to file access or code execution. Claude Mythos Preview managed 157 of 898 instances, with model safeguards disabled.
Reality
- Evidence55
- Adoption12
- Hype gap+18
- Incentives55
- Confidence58
Anthropic held the price on Claude Opus 4.8, cut fast mode to a third of its previous cost, and handed users an effort dial. That combination is what gets agents into engineering budgets.
Reality
- Evidence30
- Adoption38
- Hype gap+34
- Incentives90
- Confidence56
New agents scan, triage and open the pull request, leaving a developer to approve. The virtual patching alongside them is the tell: fixes average 50 days, exploits can land in six hours.
Reality
- Evidence34
- Adoption12
- Hype gap+42
- Incentives84
- Confidence46
Project Glasswing scanned more than 1,000 open source projects, and Anthropic says human triage became the slow part. Most teams are staffed for discovery, not response.
Reality
- Evidence22
- Adoption20
- Hype gap+32
- Incentives58
- Confidence28