Skip to content

Security1 publisher3 min readPublished Updated

The bottleneck moved: 622 CVEs in July, and no one left to write up the fixes

A Black Hat USA 2026 keynote described AI tooling that found roughly 1,000 bugs and then stalled on reporting. Patch Tuesday volume tells the same story from the other end.

The Watch · Security desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying The bottleneck moved: 622 CVEs in July, and no one left to write up the fixes
Generated illustration

What happened

  • The US government created a vulnerability clearing house named Gold Eagle to coordinate research efforts in vulnerability discovery, mitigation and fixes.
  • Microsoft's Patch Tuesday delivered 169 CVEs in April.
  • Microsoft's Patch Tuesday delivered 118 CVEs in May.
  • June's Patch Tuesday covered 571 CVEs overall, including 208 direct Microsoft CVEs.
  • July's Patch Tuesday covered another 622 vulnerabilities, including zero-days under active exploitation.

Compiled by The WatchSomething wrong?How this is made

Why it matters

Researchers at Arizona State University told Black Hat USA 2026 that their AI-assisted bug hunting found roughly 1,000 vulnerabilities and then hit a limit that had nothing to do with the models: they could not write the reports fast enough [9][13][16]. At the same time, Microsoft's Patch Tuesday count went from 169 CVEs in April to 622 in July, and the US government has stood up a vulnerability clearing house called Gold Eagle to coordinate discovery, mitigation and fixes [2][5][1].

Take the four-month curve first, because it is the part that lands in change windows. April brought 169 CVEs, May 118, June 571 overall including 208 direct Microsoft CVEs, and July 622 including zero-days under active exploitation [2][3][4][5]. June to July is an increase of 51, about 9 percent [6]. April to July is roughly a 3.7-fold increase, 453 more CVEs in a single monthly cycle [7]. The June figure is worth splitting: 363 of those 571 were not direct Microsoft CVEs [8], which means a large share of the queue arrives from third parties on someone else's schedule.

The supply side explains the shape. In a keynote, associate professor Yan Shoshitaishvili described work with his undergraduate students on using AI models for vulnerability discovery [9]. Their benchmark came from a June Washington Post article, cited by Shoshitaishvili, which reported that Anthropic's next-generation model Claude Mythos had found 479 vulnerabilities in the Linux kernel [10]. Earlier GPT generations had given the team around 300 flaws [11]. Adding Mythos-style workflows to three GPTs took them to about 600, and training the models on the properties of previously known vulnerabilities took them to about 1,000 [12][13]. That is roughly 2.1 times the Washington Post benchmark [14] and about 3.3 times their own starting point [15], from workflow and training changes rather than a new frontier model. The article's author notes that AI remains in a learning phase, where tweaking the model and workflow keeps surfacing more [24].

The interesting failure is the reporting one. In the team's usage, reporting means detailed research and a proposed fix, not just flagging the issue, and discovery outran that work [16]. The scale, in their view, calls the whole responsible disclosure process into question, which they already considered broken because disclosure often increases risk [17]. For defenders the consequence is direct: timely patching in production was already a stress point [18], and the article's author argues that growth on this curve pushes teams toward one of two bad outcomes, more unpatched software and more opportunity for criminals, or patching without testing and the compatibility breakage that follows [19].

The optimistic reading in the same piece is that discovery peaks. Human research is resource-intensive and has produced a rising but steady stream, driven by more researchers, more software and bug bounty money [21]; AI-assisted work could plausibly exhaust the back catalogue of three decades of software that no human effort could ever fully cover, after which new findings depend on model improvements [22]. The author also suggests development teams will use the same tooling to strip flaws before release [23]. That is a forecast, not a plan, and it does not help the August maintenance window.

Watch whether Gold Eagle publishes anything measurable about intake and turnaround, rather than only its remit [1]. Watch the third-party share of monthly patch counts, since that is the portion you cannot negotiate with [8]. And watch whether any disclosure programme formally changes what a report must contain, given the ASU team's finding that fix-quality write-ups are the constraint [16][17].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories