Security1 distinct publisher3 min readUpdated
CrowdStrike argues AI security evaluations have converged on vulnerability discovery because it is easy to score, while most breaches still start with stolen credentials, phishing and trusted access.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
CrowdStrike published an argument this week that AI security benchmarks have narrowed onto vulnerability discovery, exploit generation and automated proof-of-concept work, and that the narrowing is a measurement artifact rather than a reflection of defender need [1]. The consequence for anyone buying against those benchmarks is that a top score certifies competence at one initial access vector out of several.
The reason for the convergence is not mysterious. Those tasks produce binary outcomes: a vulnerability either exists or it does not, which makes them clean instruments for tracking model progress [5]. Grading is cheap, results are comparable, and leaderboards follow.
The number that matters is where the breaches actually start. Citing the Verizon 2026 Data Breach Investigations Report, CrowdStrike says vulnerability exploitation is now the single most common initial access vector, at 31% of breaches in the reporting dataset [2]. That is a real and, by CrowdStrike's characterisation, growing share [16]. It also leaves 69% of breaches beginning somewhere else [3], through credential abuse, phishing, social engineering, trusted relationships and other forms of access [4]. The paths that current AI benchmarks largely do not test outnumber the path they do test by more than two to one [13].
The rest of the work resists scoring for the same reason it is expensive. Detection and triage happen across large alert volumes, and speed there decides whether an adversary is contained in minutes or persists on the network [8]. Investigation means reconstructing events across endpoints, identities and cloud environments, and every hour spent reconstructing is an hour available for lateral movement and privilege escalation [9]. Detection engineering, which turns observed behaviour into durable detections, remains a specialised and often under-resourced discipline [10]. CrowdStrike's position is that a serious evaluation framework should measure whether AI helps with identity abuse detection, investigation, detection engineering, threat hunting and response across the lifecycle [6]. None of those produce clean binary outcomes, and all of them depend on grounded models of real adversary behaviour that public benchmarks lack [7].
That last point is the operative constraint on the alternative. Adversary emulation, which converts knowledge of real tooling, action sequences and telemetry artefacts into realistic activity, is what underpins detection engineering, hunting and validation [15]. CrowdStrike argues that AI red teaming built on hypothetical scenarios or synthetic data has limited value when it does not match the techniques and operational patterns seen in actual intrusions [11]. Its proposed test is whether a model can reproduce adversary behaviour faithfully enough to expose gaps in detection coverage, which it calls a harder and more consequential measure than vulnerability discovery alone [12].
Worth noting where the argument comes from: this is a post on CrowdStrike's own blog, not output from a neutral benchmark body [14], and the disciplines it wants added to the scoreboard are the ones its telemetry is positioned to ground.
What to watch is whether anyone builds it. A benchmark requiring real intrusion telemetry is hard to publish openly, which is convenient for whichever vendor owns the data. Ask suppliers quoting AI benchmark results which vector the benchmark covered, and what evidence they hold for the other 69% [3].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
CrowdStrike published a blog post arguing that the conversation about AI in cybersecurity has centered on vulnerability discovery, exploit generation and automated proof-of-concept development, and that benchmarks should expand beyond these.
These tasks produce binary outcomes; a vulnerability either exists or it does not, which makes them useful for measuring model progress and demonstrating increasingly sophisticated cybersecurity capabilities.
According to the Verizon 2026 Data Breach Investigations Report, as cited by CrowdStrike, vulnerability exploitation is now the most common initial access vector, accounting for 31% of breaches in the reporting dataset.
CrowdStrike characterises vulnerability exploitation as a meaningful and growing share of the problem and a strong reason to continue advancing AI capabilities in that area.
CrowdStrike notes that the 31% figure also means 69% of breaches begin through other paths.
Credential abuse, phishing, social engineering, trusted relationships and other forms of access remain central to the adversary playbook.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Thin: one self-published vendor essay, one secondhand statistic
The cluster contains exactly one source, a CrowdStrike blog post with no supplied publication date. Its only external anchor is a secondhand citation of the Verizon 2026 DBIR (31% / 69% split) with no link or methodology. Descriptive claims about defender workload and adversary emulation are internally coherent and uncontroversial, but the load-bearing normative claims - that public benchmarks lack grounded adversary models, that synthetic red teaming has limited value, and that emulation fidelity is a harder test than vulnerability discovery - carry no data, named benchmarks or third-party corroboration.
No adoption signal in supplied material
The source announces no product, release, benchmark artifact, dataset or deployment, and reports no usage, customer or pricing facts. There is nothing to measure adoption against, and inferring uptake for a proposed-but-unspecified evaluation framework would be guessing.
Mildly overstated: prescription outruns the evidence offered
The piece is deflationary about others' claims - it explicitly credits vulnerability exploitation as a meaningful and growing share and does not claim any AI capability for itself - which keeps the gap small. It tips positive because it asserts a benchmark deficiency and a difficulty ranking it does not demonstrate, and because it prescribes a 'comprehensive evaluation framework' while supplying no tasks, metrics or results. The framing implicitly favours evaluation grounded in proprietary real-intrusion telemetry without saying so.
High: vendor self-published argument that favours its own asset base
The sole source is CrowdStrike's corporate marketing blog. The argument's conclusion - that credible AI security evaluation must be grounded in real intrusion telemetry and span detection, investigation, hunting and emulation - maps directly onto the categories CrowdStrike sells and the proprietary data it holds, while devaluing benchmark approaches available to entrants without that data. No competing interest disclosure, no third-party validation and no external publisher appear in the cluster.
Low-moderate: argument clearly readable, underlying assertions unverified
Confidence is high that the source says what the ledger records - the text is explicit and internally consistent - but low that the substantive assertions hold, because there is one publisher, no date, no primary DBIR document and no measurements. Descriptive claims about SOC workload and adversary emulation are safe; the benchmark-gap and difficulty-ranking claims are single-vendor assertion.
build
A UDP packet is now enough: IKEEXT RCE moves from patch queue to fire drill1 distinct publisher
leadership
Your code review runs on human time. The intruder's agent does not.1 distinct publisher
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
security
Defender's SYSTEM race is back: ShieldBreak PoC says Microsoft's July fix never held6 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.