Security1 distinct publisher3 min readPublished
METR's public discovery series shows a sharp 2026 slope change in vulnerability reports, no comparable change across seven optimization benchmarks, and slower growth in the exploited-bug catalogues.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
A slope change in disclosure counts is a statement about what enters queues, not about what is happening to unpatched machines. That is the difficulty with the vulnerability result [2]. The series that moved hardest are the ones counting reports, and the series counting confirmed exploitation moved considerably less [5]. If you staff an advisory triage rota, the first number is the one that bills you. If you want to argue that the gap between disclosure and exploitation is closing, you need the second number, and METR's read of the CISA and VulnCheck catalogues does not hand it over [5].
The attribution inside the rise is messier than the headline suggests. On cURL and OpenSSL, most of the additional 2026 disclosures carry an AI marking [3]. On Firefox, Microsoft, the US NVD and OSV, AI credits are a small share of the increase [4]. So of the six vulnerability sources METR names, two have a legible mechanism and four rose for reasons the credit field does not explain [1]. Severity cuts the same way: METR reports that higher-severity categories generally accelerate less, though they still accelerate [6]. The growth is weighted toward the end of the distribution that consumes reviewer hours without producing an emergency.
What makes the finding worth anything is the control groups sitting next to it. METR assembled seven algorithmic-efficiency histories, including Gurobi mixed-integer programming, MIPLIB, Stockfish and the matrix-multiplication exponent [7]. Two of them, nanoGPT and CIFAR-10, already contain LLM-driven records, and none of the seven bends the way the vulnerability series does [8]. Mathematics sits in between: three named open problems from pre-existing lists were solved with AI help in 2026 [9], arXiv submissions have doubled in some areas in under a year [10], and METR itself calls the open-problem evidence weak and the domain hard to measure objectively [11]. One of the three domains produced a clear slope change [2]. A vendor claiming general acceleration now has to explain why the optimization benchmarks with dense histories and actual LLM contributions did not move.
The measurement deserves its own discount. METR says the data collection and analysis were performed by agents, that it audited the output, and that mistakes likely remain; the underlying material is in a repository built to be corrected and updated [12]. It also covers only public discoveries, and METR allows that labs may be making undisclosed ones internally [13]. Treat it as the best available public baseline rather than an instrument you would use to size a specific project's exposure.
For a defender, the useful shape of this is narrow and unglamorous. More items arrive in the same funnel, most of them below the severity line that justifies interrupting a release, and they arrive faster than any evidence that attackers have shortened their side of the clock [5][6]. That is a capacity problem in advisory handling and patch scheduling, and it is a different budget line from the one you would open if exploitation tempo had moved. The two get conflated constantly. This is the first public series that lets you keep them apart.
Ranked by verification strength, evidence, and original report placement.
METR reports the rate of vulnerabilities reported across many projects dramatically accelerated in 2026 compared with 2025, both for specific projects (cURL, OpenSSL, Firefox, Microsoft) and for aggregate databases (the US NVD and OSV).
Two of the seven optimization series, nanoGPT and CIFAR-10, include LLM-driven contributions to the plotted records, but none show a clear change in slope comparable to the changes in vulnerability or mathematical discovery.
METR describes discovery of math results as having accelerated somewhat but harder to objectively measure, and calls the open-problem solution rate weak evidence of acceleration.
A METR note titled "Have We Seen an Acceleration in Discoveries?" looked for slope changes in various metrics of discovery, highlighting January 2026 as a potential breakpoint at which AI effects might become observable.
On cURL and OpenSSL, most of the extra 2026 disclosures are AI-marked.
On Firefox, Microsoft and the aggregate databases, AI credits are a small share of the rise in disclosures.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-party quantitative series, self-audited, uncorroborated
The cluster rests on one first-party analysis that names its data sources, publishes its repository, and separates strong from weak findings. That is better than an announcement, but the source itself says data collection and analysis were done entirely by agents with mistakes likely remaining, no absolute counts or baselines appear in the text, and no second publisher or independent replication is present in the cluster.
Concentrated in two projects; flat where records are densest
There is concrete evidence of AI being used in discovery at scale in narrow slices — AI-marked reports dominate the 2026 increase on cURL and OpenSSL, and three list problems were solved with AI — but on Firefox, Microsoft, NVD and OSV the AI-credited share of the rise is small, exploited-bug catalogues lag, and seven dense optimization series show no change. Aggregate AI-attributable discovery adoption is therefore real but bounded.
Slightly overstated by attribution, not by the source
The headline framing links a 2026 discovery surge to AI, yet AI credits dominate the rise in only two of six named vulnerability sources, exploited-vulnerability growth is much slower, higher-severity acceleration is weaker, and the optimization null result runs against the loudest current claims about LLM-driven optimization. The gap is small and mostly external: METR itself labels the math evidence weak, flags competing spend-driven explanations, and publishes the null finding prominently.
Self-published mission-adjacent research, methods disclosed
This is an organisation publishing its own analysis on a topic central to its remit — measuring AI capability trajectories and the imminence of recursive self-improvement — which creates a mild pull toward finding detectable acceleration. Offsetting that, the note publishes a null result, calls its own strongest-sounding domain weakly evidenced, discloses that agents produced the data, and opens its repository. The supplied material contains no funding, commercial or product interest disclosures, so scoring stays low and structural.
Moderate
Directional findings are credible because they come from the analysing organisation with methods and repository disclosed, and the internal qualifiers (severity, KEV, AI-credit share) are self-consistent. Confidence is held at mid-range by the single-publisher cluster, the absence of independent replication, agent-generated data series with acknowledged likely errors, and the lack of absolute counts in the supplied text.
build
Fabricated SQLite CVEs cleared NVD, CISA ADP and Red Hat before anyone ran the code1 distinct publisher
invest
Forty percent, annualised: the US-Europe AI capex gap is real, the bust risk sits elsewhere2 distinct publishers
build
Firmware CVE intake: the finding is almost never a zero-day, it is a five-year-old BusyBox1 distinct publisher
build
The saved-hour fallacy: a model in which better AI makes each paper worse1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026