Build1 distinct publisher3 min readUpdated
Project Glasswing scanned more than 1,000 open source projects, and Anthropic says human triage became the slow part. Most teams are staffed for discovery, not response.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Anthropic's Project Glasswing update, as reported by dev.to, says Claude Mythos Preview and its partners found more than 10,000 high- or critical-severity vulnerabilities across major software systems [1]. In open source alone, the company says it scanned more than 1,000 projects and surfaced thousands of serious findings, with human triage becoming the slow part [2].
That last clause is the whole story. For years the constraint on security work was discovery: could someone find the bug, reproduce it, and afford enough expert review to catch the important ones first [3]. The old rhythm assumed that scarcity. A bug was found, a report was filed, a team reproduced it, someone argued about severity, someone wrote a patch, users eventually upgraded [5]. It was never fast, but it ran at roughly the speed of human discovery [5].
You do not have to accept every number to see the direction, as the dev.to writer notes [4]. Take the open source figure at its most conservative reading and it still describes at least two serious findings per project across more than a thousand projects [14]. Each of those projects has its own owner, its own release process, and its own idea of what "critical" means, which means the triage load lands on at least a thousand separate decision-makers rather than one queue [15].
The scarce skill is now the system around the finding: deciding which reports are real, prioritizing the ones that matter, patching without breaking production, shipping before attackers reach the same conclusion, and keeping maintainers from drowning in low-quality reports [6]. None of that is detection work. It is response capacity, and it is the part nobody has been hiring for.
The reason this is not merely an inbox problem: a vulnerability that sits untriaged for three weeks is not safer because an AI found it, and may be riskier, because the same class of model will make the exploitation path cheaper for everyone else [7]. The patch window stops being an operational detail and becomes part of the product [7].
The tempting version of this story is run the model, get the report, apply the patch. Real systems do not behave that way. The obvious fix breaks an integration, the technically correct fix creates a migration problem, the most severe-looking finding may be unreachable in production while the boring one in a forgotten admin path is exposed to the internet [8]. Models can write a first patch, generate a regression test, and compare similar code paths looking for variants [9]. The manual work shrinks and the judgment work expands [10].
The gap is structural: discovery is improving faster than response [11]. A large company can assign security engineers, rotate incident response, and fund tooling; a maintainer of a popular library may be doing all of it after hours, unpaid, between feature requests [12]. Dumping hundreds of generated reports into that inbox does not make the ecosystem safer, and can make it worse unless the reports are reproducible, prioritized, and paired with patches that are easy to review [13].
Watch what gets published next. Findings counts are the easy metric; the numbers that matter are time-to-triage, false-positive rate, and how many of those thousands of open source findings actually landed as merged fixes [2][13].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The author argues the scarce skill becomes the system around the finding: telling which reports are real, prioritizing the ones that matter, patching without breaking production, shipping fixes before attackers learn the same thing, and keeping maintainers from drowning in low-quality reports. That is response capacity, not detection.
In AI-assisted building the manual work shrinks but the judgment work expands, and someone still has to own the decision.
For years, security work was constrained by discovery: whether someone could find the vulnerability, reproduce it, build an exploit, and whether a small team could afford enough expert review to catch important issues before attackers did.
The author writes that you do not have to take every number at face value to see the shape of the shift.
The familiar security rhythm was: a bug was found, a report filed, a team reproduced it, someone argued about severity, someone wrote a patch, users eventually upgraded; the process was never fast enough but mostly matched the speed of human discovery.
Real systems involve tradeoffs: the obvious fix can break an integration, the technically correct fix can create a migration problem, and the most severe-looking vulnerability might be unreachable in production while a boring one in a forgotten admin path is exposed to the internet.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One secondary commentary, no primary artifact
The cluster is a single dev.to opinion post. Its only quantitative content is a two-sentence relay of an Anthropic Project Glasswing update with no link, methodology, scan scope, or false-positive data, and the author explicitly declines to stand behind the figures. Every other assertion — response capacity as the new scarcity, the widening discovery/response gap, agent patch-and-test capability — is unquantified reasoning with no independent corroboration in the cluster.
One vendor-reported program, no independent uptake
The only adoption signal is the relayed disclosure that a vendor program scanned more than 1,000 open source projects and produced thousands of findings — real scale if accurate, but reported by the vendor through a secondary source. There is no evidence in the cluster of maintainers or security teams adopting the resulting workflow, no triage or patch throughput, no named projects, and no downstream commitments, so measured adoption stays near the floor.
Big numbers, thin verification, deflationary framing
The story is framed around a striking count of critical bugs that the cluster cannot verify, and the derived per-project and per-decision-maker figures inherit that weakness — that pushes the gap positive. It is only moderately positive because the piece is not a capability advertisement: it hedges the numbers, rejects the 'run the model, apply the patch, done' story, and its main claim is a constraint (human throughput) rather than a breakthrough.
Vendor-sourced figures, independent commentator
The quantitative core originates with Anthropic describing results from its own model and program, which carries a clear promotional interest, and it reaches readers through an intermediary that did not verify it. Offsetting this, the publishing author is an individual developer with no disclosed commercial stake, no product is being sold in the post, and the argument cuts against easy AI-security claims. No sponsorship, affiliation, or funding disclosure is present either way.
Low — single unverified source
Confidence is limited by the cluster's structure: one publisher, one item, one relayed vendor datapoint, and a claim set dominated by unquantified reasoning. The publisher's emphasis and the internal argument are clear enough to read reliably, but nothing here can be triangulated, so both the figures and the gap thesis remain unconfirmed.
build
Claude's system prompt grew ninefold in two years. Version yours like code.1 distinct publisher
product
Harness hands vulnerability triage to agents, and concedes code fixes cannot keep pace2 distinct publishers
build
Count invalid JSON as a failed classification, and model choice becomes a reliability problem1 distinct publisher
security
The disclosure pipeline is triaging itself: 20,700 new CVEs, 10% more exploitation1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 17, 2026