KaliBench, 8,504 query-command pairs across 1,642 Kali tools, found no open-weight model among 24 configurations topped 42% exact-command accuracy. Scores rise when the model is handed the tool name, but production analysts describe only intent.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40
Kevin Mandia's Armadin raised $255.5 million at a valuation above $2.5 billion, bringing its total to $445 million seven months after launch. Its case against periodic pen tests rests on performance data from an exercise Armadin ran with a partner.
Perspective Coverage
6 publishers
- Builder
- Builder 24%
- Operator
- Operator 38%
- Investor
- Investor 38%
Reality
- Evidence55
- Adoption30
- Hype gap+40
- Incentives75
- Confidence60
Armadin's $255.5 million Series B and Jeeves's $110 million raise took 68% of fintech's $541 million week, by FinTech Global's count. The prior week's top two took a similar share of $1.2 billion, so the fall hit large and small rounds alike.
Reality
- Evidence45
- Adoption20
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
Armadin, the AI attack-testing startup run by Mandiant founder Kevin Mandia, raised $255.5 million at a valuation above $2.5 billion. Investors are betting that always-on AI agents will take over the budget companies now spend on periodic penetration tests.
Perspective Coverage
6 publishers
- Builder
- Builder 31%
- Operator
- Operator 32%
- Investor
- Investor 37%
Reality
- Evidence55
- Adoption20
- Hype gap+40
- Incentives80
- Confidence60
GitHub's volume counters and Chainguard's account of agent-written code both point at dependency review as the step nobody is doing. Mandiant estimates mean time-to-exploit at minus seven days in 2025.
Reality
- Evidence43
- Adoption55
- Hype gap+27
- Incentives76
- Confidence42
Anthropic says Claude Mythos Preview found and exploited previously unknown flaws in every major operating system and web browser during a month of testing. Its disclosure process keeps the rest unnamed until patches ship.
Publishers:red.anthropic.com
Reality
- Evidence36
- Adoption20
- Hype gap+35
- Incentives78
- Confidence55
The company says the architecture around a bug matters more than how fast the patch ships, and it is making that case on a stack built from its own products after watching an AI assistant fix bugs and break their dependencies.
Reality
- Evidence34
- Adoption25
- Hype gap+18
- Incentives82
- Confidence55
Fast Company reports that Treasury secretary Scott Bessent backed a 90-day government preview of new frontier models so agencies could patch first. The order that emerged asks for 30 days, voluntarily, and deployers read whatever the vendor chooses to publish.
Reality
- Evidence40
- Adoption20
- Hype gap+22
- Incentives78
- Confidence45
The June 12 order covered only foreign nationals, but Anthropic said it could not check nationality in real time, so it suspended Fable 5 and Mythos 5 for every user. Fable 5 came back 19 days later, on new usage terms.
Reality
- Evidence45
- Adoption45
- Hype gap+20
- Incentives80
- Confidence50
Anthropic's red team says the scarce reverse-engineering skill that used to give defenders weeks is no longer the bottleneck. It measured that on Firefox and Windows kernel fixes, where a 19-day median gap counts as fast.
Reality
- Evidence55
- Adoption20
- Hype gap+22
- Incentives72
- Confidence48
DeAlignAI's downloadable FP8 build is the license working exactly as written, while its self-reported 320-of-320 HarmBench run remains unchecked by any outside researcher and measures compliance, not capability.
Reality
- Evidence46
- Adoption20
- Hype gap+18
- Incentives72
- Confidence55
The Hacker News argues frontier models turn vulnerabilities into exploits at machine speed. Its remedy reorders the backlog; it does not add remediation capacity, and the text counts neither.
Reality
- Evidence14
- Adoption
- Insufficient
- Hype gap+58
- Incentives55
- Confidence52
Anthropic's model reasons its way to new bugs instead of matching known CVEs, and one vendor-sourced breakdown puts comparable capability in wider hands inside six to 24 months.
Reality
- Evidence20
- Adoption15
- Hype gap+45
- Incentives85
- Confidence38
Tenable says context can cut remediation to 1.6% of findings. If that holds, the scarce resource next year is asset truth, not patching speed.
Reality
- Evidence26
- Adoption14
- Hype gap+58
- Incentives86
- Confidence62