Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Confirmed Claude-found open source bugs outnumber upstream patches 11 to 1

Anthropic's Claude flagged more than 29,000 candidate open source vulnerabilities in six months, and upstream projects have patched 516 of them. Maintainers praise the reports, but each fix still ships only as fast as the project that tests and releases it.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Confirmed Claude-found open source bugs outnumber upstream patches 11 to 1
Generated illustration
Outside firms confirmed 5,674 of 6,123 findings reviewed Findings that outside security firms reviewed, and how many of those they confirmed as valid, as of October 2, per Anthropic's disclosure dashboard.

Bar chart of two counts from Anthropic's disclosure dashboard as of October 2: outside security firms reviewed 6,123 findings and confirmed 5,674 of them as valid.

Findings reviewed by outside security firms, as of October 2 In findings

Outside firms confirmed 5,674 of 6,123 findings reviewed (Findings reviewed by outside security firms, as of October 2)
ItemValueClaim
Reviewed by outside firms6,123 findings1
Confirmed valid5,674 findings1

What happened

  • As of October 2, the six outside security firms had confirmed 5,674 of the 6,123 findings they checked, and roughly 23,000 candidates had not been reviewed.
  • Anthropic has sent 6,157 findings to maintainers, and 584 CVE and GitHub Security Advisory identifiers have been issued, some covering the same finding.
  • Early recipients kept asking for everything else it had, Anthropic says, and nearly 5,000 unvalidated reports have gone to maintainers who requested them.
  • wolfSSL founder Todd Ouska said 72 of the 74 reports his team received were valid and five of them became CVEs.
  • curl founder Daniel Stenberg, who shut his project's bug bounty in January after a flood of AI slop, says the scanner found one of curl's worst reported vulnerabilities in years.

Why it matters

  • exposure Downstream teams are protected only once a fix ships upstream, so their window of exposure follows each project's release pace and has little to do with how fast Claude finds bugs.
  • decision Each enrolled project has to choose between triaging raw model output itself and waiting for Anthropic's outside firms to validate reports first.
  • cost Fast-track projects take on the validation Anthropic skips, including re-grading inflated severity ratings, and they pay for it in maintainer hours.

OSS Scanner, which Anthropic launched last week under its Cyber Mission [2], adds a second route for Claude's findings. In the existing disclosure program, six external security research firms reproduce and triage each candidate before it goes upstream [4]. On the new optional fast track, eligible projects get free periodic scans from Anthropic's top models, including Claude Mythos. That output goes to maintainers with no validation step at Anthropic [8]. The GitHub repository holds the enrollment and configuration tools but not the scanner itself [22], so the open source part of OSS Scanner is its sign-up tooling. Each project is built in an isolated VM with internet access cut before scanning begins, and maintainers can pause or opt out at any time [21].

In six months the disclosure route has reviewed about a fifth of Claude's candidates [26]. It rejects few of them: the firms confirmed 92.7% of the findings they checked [24]. The bigger drop comes after confirmation. Upstream projects have patched 9.1% of the confirmed bugs, about one patch for every 11 confirmed findings [27][29].

Apply that confirmation rate to the unreviewed remainder and you get about 21,000 more real bugs [25]. That estimate holds only if the firms checked a representative slice of the queue.

The fast track's precision rests on a smaller test. Expert penetration testers checked 97 critical and high-severity findings from an early scanner version across 48 projects [10]. Eighty-five cleared the disclosure bar, 11 of the other 12 were real bugs that duplicated known issues, and one was a false positive [10]. That puts real bugs at 99% and disclosure-grade findings at 88% [28]. The New Stack notes that the sample covers early output only. It excludes the 29,000 disclosure-program candidates and says little about lower-severity findings or behavior at scale [11]. For the rate to carry over, today's fast-track scans would need the early version's severity mix and accuracy. Anthropic's own technical announcement concedes that maintainers have flagged severity ratings set too high, along with cases where the scanner got a project's threat model wrong [12].

The reports themselves are well engineered. Anthropic says each one includes a self-contained reproducer and, when possible, a bisection showing where the bug was introduced. When the model can write a fix, a candidate patch comes with it [13]. OpenSSL Corporation director Anton Arapov said the reports, raw model output included, matched and sometimes beat what OpenSSL gets from human researchers [14]. He said a report with a working exploit attached is "basically job done for an engineer as you can verify it right away" [15]. PostgreSQL committer Noah Misch said several reports arrived with fixes the project could use "nearly as-is" [17]. He said fast-track access let PostgreSQL deal with the newest issues before any of them reached a GA release [18].

Those reports cut verification time. Testing the fix, backporting it and shipping it are still the maintainers' job, and sometimes the change will break behavior users depend on [23]. OpenSSH's maintainers made that call when they deliberately broke two features for security [23].

We think the evidence supports a patch backlog at the upstream end. It does not show the scanner making projects slower to fix anything: Misch describes fixes landing before a release [18]. The dashboard figures are a single reading dated October 2 [1], so any trend in the 516 patch count [5] will only appear in later readings.

What to watch

  • A later reading of Anthropic's disclosure dashboard showing whether the 516 upstream patch count is gaining on the 5,674 confirmed bugs.
  • Published precision figures for lower-severity or fast-track findings beyond the 97-finding early sample.
  • Whether Anthropic adds review capacity beyond its six outside firms to work through the roughly 23,000 unreviewed candidates.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption45
Hype gap+15
Incentives60
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    According to Anthropic's disclosure dashboard, as of October 2 outside security firms had confirmed 5,674 of the 6,123 findings they reviewed as valid.

  2. [2]

    Anthropic launched its OSS Scanner last week as part of its broader Cyber Mission.

    ReportedSupportedView cited source
  3. [3]

    Over six months of scanning some of the world's most widely used open source projects, Anthropic's models surfaced more than 29,000 candidate vulnerabilities.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. thenewstack.io

    1 article · October 11, 2026

    Claude found 29,000 possible bugs in open source. Only 516 have been fixed.

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories