Security1 distinct publisher3 min readPublished
Anthropic sent 8% of its model's 23,019 findings out for review and stopped there, citing a shortage of people to check them. Where independent scoring did happen, 14 of 27 severity ratings had to be moved.
The Watch · Security desk

Compiled by The WatchSomething wrong?How this is made
Two downgrades show what source code does not carry. For Temporal Server, Mythos assigned Critical and described an attacker controlling workflows across namespaces. Temporal's maintainers scored the issue 2.3, Low, because exploitation requires an attacker-controlled namespace that already holds a privileged internal credential, and the impact reaches only known workflows [10]. With MinIO, Mythos again said Critical, an outside security firm said High, and MinIO settled on Medium, because the attack needs an existing cluster root JWT and permits read access only [11]. Both preconditions are facts about a deployment rather than about the file the model read.
Of the 27 findings that received CVEs, 14 ratings moved after review, 13 of them down and one up, a flaw in the jq command-line JSON tool that was understated [12]. That is 51.9% of the batch [4].
The volume figures are worth doing in longhand. 1,900 of the 23,019 candidates went to outside reviewers, 8.3% of the pile [1]. 97 fixes landed, 0.42% [2]. Only 1,596 reports reached maintainers, so 304 candidates cleared external review and were still never sent on [5]. Of the 1,451 acknowledged, that leaves 1,354 with no upstream fix recorded as of May 22, 2026 [4][3].
Within the reviewed slice, 90.8% held up as real vulnerabilities [6]. Echo, the supply chain firm that compiled the figures, notes that the 1,900 were unlikely to be a random draw, so the number may describe how good the best candidates were and little else, and Anthropic has not published accuracy for the other 21,000 [7]. Anthropic attributes the drop-off to a shortage of people to check the work [5].
The exploitation result is the part that moves attacker economics, and three conditions travel with it. Anthropic built a benchmark from 50 previously discovered SpiderMonkey vulnerabilities shipping in Firefox 147 and gave each model five attempts per bug, 250 trials [13]. Mythos Preview produced working arbitrary code execution in 181 trials, 72.4%, plus partial register control in 29 more [14]. Claude Opus 4.6 managed two, under 1%, which is where the roughly 90-fold generational figure comes from [15]. Every trial starts from a crash someone else already found, the harness strips Firefox's browser sandbox and other defense-in-depth protections, and Anthropic designed and ran the evaluation, which no outside party has replicated [16]. Separately, Anthropic turned a known Linux kernel use-after-free into a working root exploit for under $2,000 in inference and under a day of runtime, chaining it with a second use-after-free it found in the kernel's traffic-control scheduler [17]. Finding a bug worth acting on still costs a few thousand dollars, with no guarantee the result is high or critical severity [18].
Echo's July 2026 survey of more than 80 senior US security leaders had 37% name detecting more vulnerabilities than they can remediate as the biggest barrier to supply chain security, and 11% name more detection or scanning as their next investment [19]. Anthropic reports the same constraint from the producing end: it has manually confirmed further vulnerabilities and not sent them to maintainers, because its own team and its external partners lack the capacity to review them [20].
Ranked by verification strength, evidence, and original report placement.
Of the 1,900 candidates reviewed externally, 1,726, or 90.8%, held up as real vulnerabilities.
Anthropic pointed Claude Mythos Preview at 281 open-source projects and collected 23,019 candidate vulnerabilities.
External security firms reviewed 1,900 of the 23,019 candidate vulnerabilities.
The other 21,119 candidates have not been reviewed by anyone outside Anthropic.
Maintainers received 1,596 reports and acknowledged 1,451; 97 fixes landed upstream and 88 findings became published security advisories, with counts current as of May 22, 2026.
Anthropic attributes the drop-off between candidates and reviewed findings to a shortage of people to check the work.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
security
Recorded Future's half-year data shows adversaries continuing to favor abusing legitimate tools and trusted platforms already inside the enterprise1 distinct publisher
build
Anthropic found 10,000 critical bugs. The bottleneck is now the person reading the report1 distinct publisher
build
ExploitGym grades agents on the step from crash input to working exploit1 distinct publisher
security
Mythos's method, not its zero-day count, is what breaks CVE-keyed vuln management1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Checkable arithmetic, one channel
The counts are specific and dated to May 22, and the sharpest evidence in the story is adversarial rather than promotional: Temporal scored its own bug 2.3 where the model said Critical, MinIO landed on Medium, and jq went the other way. That is a real outside check, but it covers 27 findings. Nothing tests the 21,119 candidates, and every figure reaches the reader through a single account of Echo's compilation.
Real patches, rounding-error scale
Code did move: 1,451 acknowledgements, 97 upstream fixes, 88 advisories, 27 CVEs. Against 23,019 candidates the fix rate is four-tenths of one percent, and roughly 1,354 acknowledged findings were still unpatched at the cutoff. Uptake is genuine and small, and it stalls at the point where a human has to look.
Overstated where it was tested
Wherever somebody independently scored this model's output, the output came down: 13 of 14 corrections lowered severity, and seven of eight Criticals did not survive. The 90.8% precision figure is drawn from a sample its own compiler says was probably not random. The ninety-fold exploit jump is the least inflated part of the story — but it starts from crashes others found, on a harness with Firefox's sandbox removed, in an evaluation its author has not had replicated.
Both suppliers have a position
Nearly every number originates with a party that benefits from it. Anthropic's own benchmark makes its newest model roughly ninety times better at exploitation than the one it replaces. Echo compiled the funnel and also produced the survey finding that too much detection is the industry's leading complaint — a tidy fit for a software supply chain vendor. Set against that, maintainers re-scoring bugs in their own projects have a pull toward the low end, which is worth remembering when reading Temporal's 2.3.
Internally consistent, externally unchecked
The figures reconcile, the cutoff date is stated, and the caveats are volunteered rather than extracted — that earns some trust. What holds the number down is structural: one publisher, one compiler, no second newsroom on the funnel and no replication of the benchmark. The claims about what nobody has reviewed are, by their nature, the ones nobody can verify.