BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Confirmed Claude-found open source bugs outnumber upstream patches 11 to 1
Anthropic's Claude flagged more than 29,000 candidate open source vulnerabilities in six months, and upstream projects have patched 516 of them. Maintainers praise the reports, but each fix still ships only as fast as the project that tests and releases it.
The Engineer · Build desk

Bar chart of two counts from Anthropic's disclosure dashboard as of October 2: outside security firms reviewed 6,123 findings and confirmed 5,674 of them as valid.
Findings reviewed by outside security firms, as of October 2 In findings
| Item | Value | Claim |
|---|---|---|
| Reviewed by outside firms | 6,123 findings | 1 |
| Confirmed valid | 5,674 findings | 1 |
What happened
- As of October 2, the six outside security firms had confirmed 5,674 of the 6,123 findings they checked, and roughly 23,000 candidates had not been reviewed.
- Anthropic has sent 6,157 findings to maintainers, and 584 CVE and GitHub Security Advisory identifiers have been issued, some covering the same finding.
- Early recipients kept asking for everything else it had, Anthropic says, and nearly 5,000 unvalidated reports have gone to maintainers who requested them.
- wolfSSL founder Todd Ouska said 72 of the 74 reports his team received were valid and five of them became CVEs.
- curl founder Daniel Stenberg, who shut his project's bug bounty in January after a flood of AI slop, says the scanner found one of curl's worst reported vulnerabilities in years.
Why it matters
- exposure Downstream teams are protected only once a fix ships upstream, so their window of exposure follows each project's release pace and has little to do with how fast Claude finds bugs.
- decision Each enrolled project has to choose between triaging raw model output itself and waiting for Anthropic's outside firms to validate reports first.
- cost Fast-track projects take on the validation Anthropic skips, including re-grading inflated severity ratings, and they pay for it in maintainer hours.
OSS Scanner, which Anthropic launched last week under its Cyber Mission [2], adds a second route for Claude's findings. In the existing disclosure program, six external security research firms reproduce and triage each candidate before it goes upstream [4]. On the new optional fast track, eligible projects get free periodic scans from Anthropic's top models, including Claude Mythos. That output goes to maintainers with no validation step at Anthropic [8]. The GitHub repository holds the enrollment and configuration tools but not the scanner itself [22], so the open source part of OSS Scanner is its sign-up tooling. Each project is built in an isolated VM with internet access cut before scanning begins, and maintainers can pause or opt out at any time [21].
In six months the disclosure route has reviewed about a fifth of Claude's candidates [26]. It rejects few of them: the firms confirmed 92.7% of the findings they checked [24]. The bigger drop comes after confirmation. Upstream projects have patched 9.1% of the confirmed bugs, about one patch for every 11 confirmed findings [27][29].
Apply that confirmation rate to the unreviewed remainder and you get about 21,000 more real bugs [25]. That estimate holds only if the firms checked a representative slice of the queue.
The fast track's precision rests on a smaller test. Expert penetration testers checked 97 critical and high-severity findings from an early scanner version across 48 projects [10]. Eighty-five cleared the disclosure bar, 11 of the other 12 were real bugs that duplicated known issues, and one was a false positive [10]. That puts real bugs at 99% and disclosure-grade findings at 88% [28]. The New Stack notes that the sample covers early output only. It excludes the 29,000 disclosure-program candidates and says little about lower-severity findings or behavior at scale [11]. For the rate to carry over, today's fast-track scans would need the early version's severity mix and accuracy. Anthropic's own technical announcement concedes that maintainers have flagged severity ratings set too high, along with cases where the scanner got a project's threat model wrong [12].
The reports themselves are well engineered. Anthropic says each one includes a self-contained reproducer and, when possible, a bisection showing where the bug was introduced. When the model can write a fix, a candidate patch comes with it [13]. OpenSSL Corporation director Anton Arapov said the reports, raw model output included, matched and sometimes beat what OpenSSL gets from human researchers [14]. He said a report with a working exploit attached is "basically job done for an engineer as you can verify it right away" [15]. PostgreSQL committer Noah Misch said several reports arrived with fixes the project could use "nearly as-is" [17]. He said fast-track access let PostgreSQL deal with the newest issues before any of them reached a GA release [18].
Those reports cut verification time. Testing the fix, backporting it and shipping it are still the maintainers' job, and sometimes the change will break behavior users depend on [23]. OpenSSH's maintainers made that call when they deliberately broke two features for security [23].
We think the evidence supports a patch backlog at the upstream end. It does not show the scanner making projects slower to fix anything: Misch describes fixes landing before a release [18]. The dashboard figures are a single reading dated October 2 [1], so any trend in the 516 patch count [5] will only appear in later readings.
What to watch
- A later reading of Anthropic's disclosure dashboard showing whether the 516 upstream patch count is gaining on the 5,674 confirmed bugs.
- Published precision figures for lower-severity or fast-track findings beyond the 97-finding early sample.
- Whether Anthropic adds review capacity beyond its six outside firms to work through the roughly 23,000 unreviewed candidates.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption45
- Hype gap+15
- Incentives60
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
According to Anthropic's disclosure dashboard, as of October 2 outside security firms had confirmed 5,674 of the 6,123 findings they reviewed as valid.
- [2]
Anthropic launched its OSS Scanner last week as part of its broader Cyber Mission.
- [3]
Over six months of scanning some of the world's most widely used open source projects, Anthropic's models surfaced more than 29,000 candidate vulnerabilities.
- [4]
Anthropic's human review pipeline relies on six external security research firms to reproduce and triage findings, and has worked through only about 6,000.
- [5]
As of October 2, only 516 vulnerabilities had been patched upstream.
- [6]
Anthropic sent 6,157 findings to maintainers, and 584 CVE and GitHub Security Advisory identifiers were issued, though some findings received both.
- [8]
Under the fast track, Anthropic gives eligible projects free, periodic scans powered by its top models, including Claude Mythos, and the findings go directly to maintainers without anyone at Anthropic validating them first.
- [9]
Anthropic says projects that received its first reports increasingly asked for everything else it had; it has sent nearly 5,000 unvalidated reports to maintainers who requested them and calls the arrangement an optional fast-track.
- [10]
Expert penetration testers checked 97 critical and high-severity findings that an early version of the scanner produced across 48 projects; 85 cleared the bar for disclosure, 11 of the remaining 12 were real bugs that duplicated known issues or other findings, and one was a false positive.
- [11]
The New Stack notes the 97-finding figures describe the scanner's early output, not the 29,000 candidates from Anthropic's separate disclosure program, and say little about lower-severity findings or how the scanner will perform at scale.
- [12]
In its technical announcement, Anthropic acknowledged that maintainers have reported inflated severity ratings and cases where the scanner misunderstood a project's threat model.
- [13]
Anthropic says reports include a self-contained reproducer, an explanation that when possible identifies where the bug was introduced through bisection, and a candidate patch when the model can produce one.
- [14]
Anton Arapov, director of OpenSSL Corporation, said the reports Anthropic sent, raw model output included, matched and sometimes beat what the project gets from human researchers.
- [15]
"basically job done for an engineer as you can verify it right away." (Anton Arapov, on a report with a working exploit attached)
- [16]
Todd Ouska, founder of wolfSSL, reported that 72 of the 74 reports his team received were valid and five became CVEs.
- [17]
PostgreSQL committer Noah Misch said several reports arrived with fixes the project could use "nearly as-is."
- [18]
Misch credited fast-track access with letting PostgreSQL address the newest issues before they shipped in a GA release.
- [19]
curl founder Daniel Stenberg spent last year publicly complaining about AI-generated slop hitting the project's bug bounty, which curl shut down in January.
- [20]
Stenberg now says OSS Scanner found multiple issues in curl, including one of the worst vulnerabilities the project has seen reported in the last few years.
- [21]
Anthropic builds each project in an isolated VM and cuts off internet access before scanning begins; reports go to maintainers by email, and they can pause or opt out at any time.
- [22]
The GitHub repository contains the enrollment and configuration tools, but not the scanner itself.
- [23]
Maintainers still have to test the fix, handle backports, and ship it, sometimes knowing the change will break behavior users rely on; OpenSSH's maintainers deliberately broke two features in the name of security.
- [24]
The outside firms confirmed 92.7% of the findings they reviewed.
- [25]
If the roughly 23,000 unreviewed candidates confirm at the reviewed rate, they contain about 21,000 real bugs.
- [26]
Anthropic's outside firms have reviewed about a fifth (at most 21%) of Claude's candidate vulnerabilities.
- [27]
Upstream patches cover 9.1% of the confirmed bugs.
- [28]
In the 97-finding early sample, 99% were real bugs and 88% were disclosure-grade.
- [29]
There are about 11 confirmed bugs for every upstream patch.
Sources
1 independent publisher whose own reporting we read for this story.
- thenewstack.ioClaude found 29,000 possible bugs in open source. Only 516 have been fixed.
1 article · October 11, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.