Security1 publisher3 min readPublished
Benchmarking AI On Bug Hunting Scores 31% Of The Breach Problem
CrowdStrike argues AI security evaluations have converged on vulnerability discovery because it is easy to score, while most breaches still start with stolen credentials, phishing and trusted access.
The Watch · Security desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- CrowdStrike published a blog post arguing that the conversation about AI in cybersecurity has centered on vulnerability discovery, exploit generation and automated proof-of-concept development, and that benchmarks should expand beyond these.
- These tasks produce binary outcomes; a vulnerability either exists or it does not, which makes them useful for measuring model progress and demonstrating increasingly sophisticated cybersecurity capabilities.
- According to the Verizon 2026 Data Breach Investigations Report, as cited by CrowdStrike, vulnerability exploitation is now the most common initial access vector, accounting for 31% of breaches in the reporting dataset.
- CrowdStrike characterises vulnerability exploitation as a meaningful and growing share of the problem and a strong reason to continue advancing AI capabilities in that area.
- CrowdStrike notes that the 31% figure also means 69% of breaches begin through other paths.
Compiled by The WatchSomething wrong?How this is made
Why it matters
CrowdStrike published an argument this week that AI security benchmarks have narrowed onto vulnerability discovery, exploit generation and automated proof-of-concept work, and that the narrowing is a measurement artifact rather than a reflection of defender need [1]. The consequence for anyone buying against those benchmarks is that a top score certifies competence at one initial access vector out of several.
The reason for the convergence is not mysterious. Those tasks produce binary outcomes: a vulnerability either exists or it does not, which makes them clean instruments for tracking model progress [5]. Grading is cheap, results are comparable, and leaderboards follow.
The number that matters is where the breaches actually start. Citing the Verizon 2026 Data Breach Investigations Report, CrowdStrike says vulnerability exploitation is now the single most common initial access vector, at 31% of breaches in the reporting dataset [2]. That is a real and, by CrowdStrike's characterisation, growing share [16]. It also leaves 69% of breaches beginning somewhere else [3], through credential abuse, phishing, social engineering, trusted relationships and other forms of access [4]. The paths that current AI benchmarks largely do not test outnumber the path they do test by more than two to one [13].
The rest of the work resists scoring for the same reason it is expensive. Detection and triage happen across large alert volumes, and speed there decides whether an adversary is contained in minutes or persists on the network [8]. Investigation means reconstructing events across endpoints, identities and cloud environments, and every hour spent reconstructing is an hour available for lateral movement and privilege escalation [9]. Detection engineering, which turns observed behaviour into durable detections, remains a specialised and often under-resourced discipline [10]. CrowdStrike's position is that a serious evaluation framework should measure whether AI helps with identity abuse detection, investigation, detection engineering, threat hunting and response across the lifecycle [6]. None of those produce clean binary outcomes, and all of them depend on grounded models of real adversary behaviour that public benchmarks lack [7].
That last point is the operative constraint on the alternative. Adversary emulation, which converts knowledge of real tooling, action sequences and telemetry artefacts into realistic activity, is what underpins detection engineering, hunting and validation [15]. CrowdStrike argues that AI red teaming built on hypothetical scenarios or synthetic data has limited value when it does not match the techniques and operational patterns seen in actual intrusions [11]. Its proposed test is whether a model can reproduce adversary behaviour faithfully enough to expose gaps in detection coverage, which it calls a harder and more consequential measure than vulnerability discovery alone [12].
Worth noting where the argument comes from: this is a post on CrowdStrike's own blog, not output from a neutral benchmark body [14], and the disciplines it wants added to the scoreboard are the ones its telemetry is positioned to ground.
What to watch is whether anyone builds it. A benchmark requiring real intrusion telemetry is hard to publish openly, which is convenient for whichever vendor owns the data. Ask suppliers quoting AI benchmark results which vector the benchmark covered, and what evidence they hold for the other 69% [3].