Product1 distinct publisher3 min readPublished
Booz Allen scored 18 frontier models on a live intrusion and placed Claude Sonnet 5 fifteenth, then paired it with an attack harness and watched it rival the winner. The result: anyone tiering risk by model name is reading a column that measures the wrong object.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The tab is usually called Approved Models, and somewhere to the right there is a column marked risk tier. Booz Allen's table fills that column in nicely: Grok-4.5 on 49, GPT-5.6 Sol on 46, Meta's Muse Spark 1.1 and Moonshot's Kimi K3 tied on 38, Alibaba's Qwen3-Coder last on 4 [8]. The firm's own conclusion is that the column is measuring the wrong object, because the model is no longer the unit of risk [13].
The test was built to produce exactly that column and no more. Each model sat on its own attacker machine and issued commands one at a time, with no tool menu and no supporting software, because Booz Allen wanted the model stripped of the engineering that normally surrounds it [3]. It then scored what the traffic records and intrusion-detection sensors proved rather than what the models claimed to have done [5], which is more than most benchmarks bother with. What that design cannot tell you is what someone with a laptop and a fortnight would build around any of these.
Sonnet 5 came 15th of 18 on a score of 13, and rose to rival Mythos once Booz Allen attached an attack harness, the software that wires a model to hacking tools and keeps it on task [10]. Booz Allen puts the closed gap at 67 points [11]. Run it as a ratio instead and 13 becomes roughly 80, a factor of about six [1]. Either way, the swing from adding plumbing is bigger than the 45 points that separate second place from last [2]. The variable that moved the number most is not on the leaderboard at all.
An attacker takes whatever model is available and unlikely to refuse, then spends the effort on the harness, because the harness is the part that keeps the model focused, lets it recover from failure and chains single actions into a sustained operation [12]. Security teams, by contrast, tend to reach for the top of the ranking. The refusal finding points the same way. One model declined a job for want of credentials, and its cyber-tuned sibling took the identical job and completed it [15], which is why Booz Allen treats guardrails as a function of context and configuration rather than a property of a model [16].
None of this comes out tidy, though, for two reasons. Booz Allen concedes it has not tested Chinese or open-weight models with optimised harnesses, while saying its results strongly suggest fully capable combinations already exist [14], so the cheapest dangerous pairing carries no score. And on real-world vulnerability research the report says all nine frontier models scored zero against an unseen production flaw, in the same breath as saying Anthropic's frontier models spotted it and Mythos could exploit it [18][19]. The Next Web reads the zero as an artefact of scoring [19]. Anyone banking the breathing-room framing, and Booz Allen's expectation that most of the 18 will catch up to Mythos [21], is banking on that reading rather than on the number.
The risk-tier column, then, needs two entries per row rather than one. First, what the model did alone under log-verified conditions. Second, whether anyone has measured it wrapped in the smallest harness a competent adversary would reach for. Where the second is blank, the row is unscored rather than low risk. Sonnet 5's 13 measured a bare configuration, one stripped of the harness that turns a model into an operation, and no attacker builds it that way.
Ranked by verification strength, evidence, and original report placement.
Booz Allen published the Cyber Weapon Index on Wednesday, scoring nine American and nine Chinese AI models (18 in total) on how far each could get through a real intrusion against a live corporate network with nobody steering.
Only Anthropic's Claude Mythos reached the end of the intrusion task.
Each model got its own attacker machine and issued commands one at a time; Booz Allen gave them no tool menu and no supporting software, in order to measure what a model can do alone, stripped of the engineering that normally surrounds it.
Scoring combined two halves: whether a model could find vulnerabilities in compiled software with no source code, and how far it advanced through an intrusion against a defended Active Directory network from first access to full domain admin control.
Booz Allen says it scored what the network logs proved, using traffic records, security logs and intrusion-detection sensors, rather than what the models claimed.
Claude Mythos scored 80; given a stolen employee credential it took administrator control on every attempt and worked out its own route to higher privileges rather than following a script.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
CrowdStrike will police the OpenAI agents it also puts to work1 distinct publisher
invest
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead1 distinct publisher
invest
Three bodies, one price sheet: why no single AI winner is worth betting on1 distinct publisher
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Telemetry-scored, single-sourced, self-contradicting in one place
The intrusion half is unusually well grounded for a capability claim: progress was credited against traffic records, security logs and sensors rather than model self-report, and the harness experiment was run by the same firm whose ranking it demolishes. Everything above that floor is weaker. One publisher retells one vendor document, the vulnerability-research section says all nine models scored zero and then says Anthropic's found the flaw, and the 10,000-vulnerabilities-in-May line arrives with no origin at all. Strong instrumentation, thin verification chain.
Published and productised, not yet replicated
What exists is a document, a product shipped beside it, and a rival lab's own threshold statement days earlier. What does not exist is anyone reproducing the harness result, any disclosed use of the index in procurement or defence planning, or a single measurement of the combination the report itself calls most dangerous — open-weight or Chinese models with optimised harnesses. Adoption is scored on the publication events actually visible, and they are all one firm's.
The reporting deflates, the report escalates
Two directions cancel unevenly. The Next Web's framing is corrective — it tells you the ranking is the wrong takeaway and names who profits from the fix — which is understatement relative to a story that could have run as 'AI hacks corporate network.' Pushing the other way is the report's own reach: mainstream AI-enabled attacks called imminent, most of the field expected at Mythos's level inside six months, a 95% reduction figure verified by nobody, and an assertion that fully capable open-weight combinations already exist based on tests never run. Net, the claims sit modestly ahead of what was measured.
The finding and the remedy have the same author
A consultancy measured a threat, concluded the unit of risk is the whole system rather than the model, and released a system-level defence product in the same breath. The unverified 95% figure argues for that product. The policy asks — enforceable containment deadlines for critical infrastructure, a national testing programme for foreign and open-weight models, governed offensive access for vetted defenders — describe a market for exactly the firm's services. None of this makes the harness result wrong; it does mean the alarming parts and the sellable parts point the same way, and only one publisher in our coverage says so.
Act on the harness point, wait on the timeline
The core structural finding would survive most corrections: it comes from the vendor's own experiment, cuts against its own headline ranking, and is internally consistent. Confidence stalls below that because a single retelling of a single interested document is all we have, one passage contradicts itself, and the forward-looking claims are expectations dressed as findings. Enough to stop tiering exposure by model name; not enough to plan around six months.