Build1 publisher3 min readPublished
Anthropic found 10,000 critical bugs. The bottleneck is now the person reading the report
Project Glasswing scanned more than 1,000 open source projects, and Anthropic says human triage became the slow part. Most teams are staffed for discovery, not response.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Anthropic's Project Glasswing update says Claude Mythos Preview and its partners found more than 10,000 high- or critical-severity vulnerabilities across major software systems.
- In open source alone, Anthropic says it scanned more than 1,000 projects and surfaced thousands of serious findings, with human triage becoming the slow part.
- For years, security work was constrained by discovery: whether someone could find the vulnerability, reproduce it, build an exploit, and whether a small team could afford enough expert review to catch important issues before attackers did.
- The author writes that you do not have to take every number at face value to see the shape of the shift.
- The familiar security rhythm was: a bug was found, a report filed, a team reproduced it, someone argued about severity, someone wrote a patch, users eventually upgraded; the process was never fast enough but mostly matched the speed of human discovery.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Anthropic's Project Glasswing update, as reported by dev.to, says Claude Mythos Preview and its partners found more than 10,000 high- or critical-severity vulnerabilities across major software systems [1]. In open source alone, the company says it scanned more than 1,000 projects and surfaced thousands of serious findings, with human triage becoming the slow part [2].
That last clause is the whole story. For years the constraint on security work was discovery: could someone find the bug, reproduce it, and afford enough expert review to catch the important ones first [3]. The old rhythm assumed that scarcity. A bug was found, a report was filed, a team reproduced it, someone argued about severity, someone wrote a patch, users eventually upgraded [5]. It was never fast, but it ran at roughly the speed of human discovery [5].
You do not have to accept every number to see the direction, as the dev.to writer notes [4]. Take the open source figure at its most conservative reading and it still describes at least two serious findings per project across more than a thousand projects [14]. Each of those projects has its own owner, its own release process, and its own idea of what "critical" means, which means the triage load lands on at least a thousand separate decision-makers rather than one queue [15].
The scarce skill is now the system around the finding: deciding which reports are real, prioritizing the ones that matter, patching without breaking production, shipping before attackers reach the same conclusion, and keeping maintainers from drowning in low-quality reports [6]. None of that is detection work. It is response capacity, and it is the part nobody has been hiring for.
The reason this is not merely an inbox problem: a vulnerability that sits untriaged for three weeks is not safer because an AI found it, and may be riskier, because the same class of model will make the exploitation path cheaper for everyone else [7]. The patch window stops being an operational detail and becomes part of the product [7].
The tempting version of this story is run the model, get the report, apply the patch. Real systems do not behave that way. The obvious fix breaks an integration, the technically correct fix creates a migration problem, the most severe-looking finding may be unreachable in production while the boring one in a forgotten admin path is exposed to the internet [8]. Models can write a first patch, generate a regression test, and compare similar code paths looking for variants [9]. The manual work shrinks and the judgment work expands [10].
The gap is structural: discovery is improving faster than response [11]. A large company can assign security engineers, rotate incident response, and fund tooling; a maintainer of a popular library may be doing all of it after hours, unpaid, between feature requests [12]. Dumping hundreds of generated reports into that inbox does not make the ecosystem safer, and can make it worse unless the reports are reproducible, prioritized, and paired with patches that are easy to review [13].
Watch what gets published next. Findings counts are the easy metric; the numbers that matter are time-to-triage, false-positive rate, and how many of those thousands of open source findings actually landed as merged fixes [2][13].