Build1 distinct publisher3 min readPublished
Half of those rejections say the same thing, that the proposing model pushed severity past the CVSS evidence it had just cited. That is a real finding about the proposer, and it still says nothing about whether a same-family reviewer would have caught it.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Start with the denominators, because in this write-up they carry the argument. Ninety-one rejections across 139 verdicts leaves 48 verdicts that ended in ratification [7][1], and 91 divided by 139 is 65.5% [2]. A ratified verdict closes a finding. A rejected one buys the finding exactly one re-proposal with the feedback attached [9]. That makes the accounting closed enough to solve: if F findings reached a decision and k of them were ratified on the second pass, the two published counts force 2F + k = 187 [3].
Set the ratification share at exactly half and you need F = 96, which would require 96 rejections, and there were 91 [4]. The nearest fits that respect both counts are 93 findings with a single successful re-proposal, at a 52% ratification share, or 88 findings with 11 successful ones at 55% [5]. So between roughly 2% and 22% of re-proposals survive the second look [5]. The feedback loop is doing very little rescuing.
The adoption cost falls out of the same numbers. 139 verdicts implies 139 proposals and 139 reviews, so 278 model calls against at most 93 findings, about three calls per finding before anybody receives a ticket [6]. That is the sticker price of challenging every decision before it becomes state [2].
The severity split needs more care than the post gives it. 24 critical against 24 high is 50%, not the 55% reported [17][7]; the ratified pair, 22 against 32, does come out at 41% [18][8]. The subtotals also do not reconcile with the top line, 48 and 54 against 91 rejections and 48 ratifications [9]. The direction is probably right, but I would want the export before quoting the percentage.
What the reviewer actually catches is narrow, which is why it catches anything. Both quoted rejections compare a proposed severity against evidence the proposal itself cited: an escalation to critical from a CVSS base of 7.8 [13], and a critical rating set against a scanner CVSS of 5.4 and an NVD description at high, 7.8 [14]. That is text checked against text. For the 65% to transfer to your pipeline, proposals have to carry their evidence inline, the reviewer has to see the same evidence, and the rubric has to make escalation past the citation a rejectable condition [3][4]. A reviewer grading judgment with no citations to check would emit a rate that means nothing.
The reviewer's own accuracy sits outside the measurement. The author's conclusion that the triage model inflates severity rests on the reviewer's stated reasons and on that split [16][17][18], and no human severity call is scored against either model. Read strictly, 65% is consistent with a proposer that inflates and with a reviewer that is merely strict.
The craft worth copying is the correction. An earlier version of the piece read one silent cycle as restraint; the honest version is that both workers woke at 09:01 UTC on 28 August, on cycle 30693, with an empty acceptance sweep and an empty set of SLA clocks [21][22]. Actual deliberation shows up elsewhere, in cycle 9003's wait=9, nine findings evaluated and nine left alone because their deadlines had not arrived [23]. The duplicate-delivery guard had to be provoked by hand, since Pub/Sub delivers at least once and production had declined to oblige [25].
Ranked by verification strength, evidence, and original report placement.
Half of all rejections cite the same reason.
One rejection reads verbatim: "The severity is escalated to critical without evidence supporting such a jump from the CVSS base of 7.8, and the remediation is vague."
Another rejection reads verbatim that the severity is rated critical despite the scanner's CVSS being 5.4 and the NVD description indicating a high (7.8) severity, creating a mismatch between evidence and proposal.
A single rejection can cite more than one reason category, so the category shares total more than 100%.
The author concludes the triage model has a consistent bias toward inflating severity past its own cited evidence.
The author spent five days building an autonomous system that owns the vulnerability remediation lifecycle, the six weeks after a scan.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Giving one deterministic layer the only write privilege demotes the model to a proposer1 distinct publisher
build
Keycloak's forgot-password flow hands over admin accounts, and the fix is a same-day call1 distinct publisher
build
Firmware CVE intake: the finding is almost never a zero-day, it is a five-year-old BusyBox1 distinct publisher
build
The demo passed because Cloud Run didn't scale: a correlation bug that emits no error1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-party, granular, not reconciled
Everything traces to one dev.to post written by the person who built the system, and its numbers do not close: an exactly even ratification share is unreachable from 91 rejections across 139 verdicts, and the 24-to-24 rejected split is 50% critical rather than the 55% printed beside it. What lifts it above anecdote is texture — rejection reasons quoted verbatim, cycle numbers and log lines reproduced, and an explicit rewrite of an earlier, tidier version of the idle morning.
One builder, five days in
The install base is the author's laptop and one Google Cloud project. All usage on record is his own: 139 verdicts by 31 August, a scheduled loop that woke unattended on 28 August, and a redelivery test he had to trigger himself because production never produced one. No second operator, no external review of the logs.
Careful conclusions, loose arithmetic
The headline conclusion is the rare one that undersells: the author refuses to claim a same-family reviewer would have missed the bias and names the experiment he skipped. The overstatement sits lower down, in the figures. A 50% critical split is printed as 55% and described as confirmation, and the severity subtotals never meet the top-line verdicts — so the numeric support for the bias reads tidier than it is.
No sponsor, reputational upside
No vendor, no funding round, nothing for sale — but a five-day solo build write-up pays in credibility, and the version where the architecture looks clever is the profitable one. Two visible counterweights: he prints the unflattering denominator next to the flattering one, and he rewrites his own account of the idle cycle to say the agent's restraint was trivial rather than considered.
Single voice, small n
One publisher, one participant, 139 verdicts on a system less than a week old, and internal figures that do not reconcile. The architectural facts — which models, which topology, which log lines — are credible because they are specific and checkable in principle. The generalisable claim, that cross-family review catches what self-review would not, has no support at all beyond one uncontrolled run.