Published Security3 min read
XBOW says the offensive-AI problem is the cheap seats, not the frontier
New testing from XBOW argues that open-weight and mid-tier models have crossed a competence threshold, and that their low cost lets an attacker run them again and again until they win.
Not a builder's beat, but builders have a standing stake in it.See today for builders

What happened
- Research from XBOW published this week shows a growing class of both proprietary and open-source models becoming strategically important in the offensive security ecosystem, with researchers warning the industry's 'middle class' of smaller models may pose a greater threat over the long term than frontier models.
- Albert Ziegler, head of AI at XBOW: "It's not even that the open-source variants or...not quite frontline competitors are catching up [to frontier models] as such. It's that they are crossing a certain threshold, which means that suddenly they are providing net value at a cheaper price."
- Models including Z.ai's open-weight GLM-5.2, xAI's Grok 4.5, Anthropic's Opus 4.7 and Meta's Muse Spark 1.1 still perform very strongly at many hacking and exploitation tasks that worry policymakers.
- As recently as six months ago, testing on mid-tier class models showed they struggled to complete 'moderately complex' agentic tasks; today's middle class largely can.
- The relative cheapness of mid-tier models means users can spend many times more resources, running them repeatedly, to solve the same challenges.
Compiled by The WatchSomething wrong?How this is made
Why it matters
XBOW published testing this week arguing that the models most relevant to offensive security are not the most capable ones but the cheap ones that recently became good enough [1]. According to Albert Ziegler, the company's head of AI, the change is not that open-weight and near-frontier models are catching up to frontier systems as such, but that they have crossed a threshold where they suddenly provide net value at a cheaper price [2].
The models XBOW puts in this class include Z.ai's open-weight GLM-5.2, xAI's Grok 4.5, Anthropic's Opus 4.7 and Meta's Muse Spark 1.1, all of which the company says still perform very strongly on many of the hacking and exploitation tasks that worry policymakers [3]. Six months ago, XBOW's testing showed this tier struggled to complete moderately complex agentic tasks; today it largely can [4]. That is the entire argument, and it is an economic one: because the models are cheap, a user can spend many times more resources running them repeatedly against the same challenge [5]. Ziegler's version is blunter. Cheaper models can be given more time, and they "come from behind and leapfrog the big frontier model," which did not work half a year ago because a long-horizon agentic run on an open-source model "would just get lost" [6].
The near-frontier tier is moving too. XBOW recorded one of its best-ever exploitation benchmark performances from GPT 5.5 [7], and the report calls the gap between GPT 5 and 5.5 one of the clearest 2026 leaps in autonomous web application testing [8]. The number that matters is the miss rate: 10 percent for GPT 5.5 against 40 percent for GPT 5 [9], a 30 point spread [10]. More usefully for defenders, GPT 5.5 scored higher in tests without source code access, while GPT 5 leaned heavily on the code [11]. XBOW's read is that working without the code, as an attacker would, the newer model beat a version that could read it, and that what produced findings was reaching and proving a vulnerability against the running system rather than inferring it from a pattern in source [12]. Live interaction with the target mattered more than code access [13].
The cost side of the ledger comes from Anthropic, which tested Mythos Preview, used in Project Glasswing, and Opus 4.8 on how fast multi-agent swarms could find bugs across 15 open-source projects [14]. Agents working individually against assigned core directories found 21 vulnerabilities; the coordinating swarm found 266 [15], roughly 12.7 times as many [16]. The runs consumed 6.5 million and 27 million tokens [17], and Anthropic notes that few individuals or organisations can underwrite that [18]. Frontier models such as Mythos and GPT 5.6 are more capable on individual security tasks, but at exponentially higher token cost [19]. That is the trade the middle class exploits.
One caveat on the swarm result: the coordination was not human-like. Anthropic's separate video game experiment saw Opus 4.6 fail to coordinate and produce bad results, while Mythos and Opus 4.8 got better results by barely coordinating at all, siloing themselves and largely failing to merge their work [20]. Anthropic also observes that agents are more homogeneous than humans [21].
What to watch: whether policy attention aimed at frontier model gating survives the finding that an open-weight model plus patience does comparable work, and whether XBOW's benchmarks start reporting cost per confirmed finding rather than capability per task. Also watch the miss rate as a defensive metric. A 10 percent miss rate on a cheap model run five times is a different threat model than one expensive pass.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Research from XBOW published this week shows a growing class of both proprietary and open-source models becoming strategically important in the offensive security ecosystem, with researchers warning the industry's 'middle class' of smaller models may pose a greater threat over the long term than frontier models.
- [2]
Albert Ziegler, head of AI at XBOW: "It's not even that the open-source variants or...not quite frontline competitors are catching up [to frontier models] as such. It's that they are crossing a certain threshold, which means that suddenly they are providing net value at a cheaper price."
- [3]
Models including Z.ai's open-weight GLM-5.2, xAI's Grok 4.5, Anthropic's Opus 4.7 and Meta's Muse Spark 1.1 still perform very strongly at many hacking and exploitation tasks that worry policymakers.
- [4]
As recently as six months ago, testing on mid-tier class models showed they struggled to complete 'moderately complex' agentic tasks; today's middle class largely can.
- [5]
The relative cheapness of mid-tier models means users can spend many times more resources, running them repeatedly, to solve the same challenges.
- [6]
Ziegler: "Because these models are cheaper, it's okay to give them more time, and they come from behind and leapfrog the big frontier model. Now, that didn't work half a year ago because...if you wanted to run some open-source model on a complex task in an agentic way...on a long horizon, then it would just get lost."
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- cyberscoop.comdjohnsonAug 13AI’s ‘middle class’ has gotten dramatically better at hacking
Additional citations
- XBOW research, reported by CyberScoop
- Albert Ziegler, XBOW
- XBOW research
- XBOW
- XBOW report
- Anthropic
- CyberScoop, citing Anthropic research
- XBOW / CyberScoop
- Anthropic blog



