Three of the Crosswork flaws score a flat 10.0, and Cisco says each CVE bundles several underlying defects. No exploitation reported in the wild so far.
Perspective Coverage
3 publishers
- Builder
- Builder 28%
- Operator
- Operator 62%
- Investor
- Investor 10%
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence70
AWS has published a 14,822-sample set that pairs real vulnerability patterns with controls that stop them. With direct prompting the 12 models it scored flagged between 41% and 99% of safe samples as exploitable.
Reality
- Evidence58
- Adoption18
- Hype gap+15
- Incentives65
- Confidence57
Its review of fiscal 2024 and 2025 says opportunistic scanning of known, internet-exposed flaws drove most compromises. The fix, it argues, belongs to software producers, not to defenders patching faster.
Reality
- Evidence58
- Adoption18
- Hype gap+15
- Incentives55
- Confidence55
A developer ran 200 flagged snippets past two frontier models with identical prompts. One cleared 51% of the false alarms; the other cleared 20% and agreed with 90% of what it saw. The countermeasures are a model property.
Reality
- Evidence24
- Adoption9
- Hype gap+32
- Incentives46
- Confidence33
PlannerCritic's author tried to inject his own engine. The blocks arrived as feasibility verdicts rather than safety strings, which is a design worth copying and a limit worth reading closely.
Reality
- Evidence44
- Adoption12
- Hype gap+24
- Incentives72
- Confidence41
The hosted Claude Security beta bills as ordinary Claude usage with no platform fee, while Mythos 5 stays with vetted partners. What runs the Enterprise scans is not settled by the sources.
Reality
- Evidence38
- Adoption33
- Hype gap+22
- Incentives72
- Confidence40
GitHub's August release retunes Actions queries for cache poisoning, output clobbering and untrusted checkouts. Most of the work went into cutting false positives, not adding coverage.
Reality
- Evidence58
- Adoption28
- Hype gap+6
- Incentives44
- Confidence60
A Secure Code Warrior and RMIT study of six frontier models across 11 frameworks found no universal winner and no link between token cost and secure output.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+20
- Incentives68
- Confidence40