Build1 distinct publisher3 min readPublished
Kellogg's 2017 review found that only 106 of 302 investigations recommended anything at all. The ones that did averaged close to seven items, mostly training and policy reminders.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The arithmetic inside the 106 is the part worth an hour. Those reviews that recommended anything recommended 731 things in total [5], which works out to about 6.9 items each [9]. That is not a gentle slope from thorough to thin. An investigation either ended with a pile of action items or it ended with an empty recommendation section, and 196 of the 302 ended empty [3], roughly 65 percent of the set [4].
Then look at what the 731 were made of. Training at 20 percent, process change at 19.6 percent and policy reinforcement at 15.2 percent [6] come to 54.8 percent of the total [7], or about 401 individual items [8]. Kellogg's team judged that the most frequently recommended solutions were the weaker interventions, the ones less likely to prevent recurrence [10]. The recurrence duly arrived: multiple event types repeated during the study period despite repeated analyses [11].
So there are two failure modes stacked on each other, and they are not the same problem. One is a remedy-strength problem inside the minority of reviews that produced remedies. The other is a much larger completion problem, where a structured investigation into real harm [17] ran to the end and named nothing anyone had to do. Review technique addresses neither. The dev.to essay is straight about the limit of the evidence here: nobody has measured that better post-mortems produce worse follow-through, and what the patient-safety literature actually documents is decoupling, with remedy strength varying independently of analysis quality [13]. A careful investigation and a bored one land on the same three outputs [14].
The essay offers a psychological candidate for why thoroughness might even compete with follow-through, resting on experimental work in which structured ritual lowers anxiety and the neural response to performance errors, and it labels that a hypothesis rather than a finding [15]. Treat it that way. The structural reading needs no psychology: if the artifact that leaves the room is a document, the document is what gets produced well, and Peerally and colleagues gave the result a name in a companion paper the same year, organisational forgetting [12].
The transfer to software is an analogy the essay draws rather than a count it supplies: the two-hour review, the accurate timeline, and the bug or its sibling back in production inside the quarter [18]. If your incident process resembles the medical one in structure, the number to instrument is not review depth. It is the share of reviews that exit with an owned, dated change to a system, and the share of those that are still open ninety days later. Kellogg's institution had eight years of documents [1] and repeat harm to go with them [11]. Only the tracker, and the engineering work it forces, gets a vote on the next failure [16].
Ranked by verification strength, evidence, and original report placement.
The study states: "In 106 RCAs, solutions were proposed."
There were 731 proposed solutions across the analyses that proposed any.
The most common proposed solution types were training (20%), process change (19.6%) and policy reinforcement (15.2%).
A 2017 study in BMJ Quality & Safety led by Kathryn Kellogg, titled "Our current approach to root cause analysis: is it contributing to our failure to improve patient safety?", reviewed 302 root cause analyses conducted over eight years at a major academic medical center.
196 of the 302 reviewed root cause analyses proposed no solution at all.
Reviews with no proposed solution accounted for about 65 percent of the sample.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed core, single-source relay
The load-bearing numbers are direct quotations from a named 2017 BMJ Quality & Safety study and a companion Peerally paper, which is unusually solid grounding for an opinion essay. But the cluster contains exactly one item, no link-level access to the primary papers, no independent corroboration, and the derived figures (196, 65%, 54.8%, ~401, ~6.9 per RCA) are arithmetic rather than reported results. The essay's own extensions — the hazard-control mapping, the ritual hypothesis, the software transfer — are explicitly unmeasured.
No adoption signal in cluster
The cluster contains no release, deployment, benchmark, pricing, licensing, security or usage-disclosure event. It is a commentary essay about a nine-year-old study; nothing in the supplied material indicates uptake of any practice, tool or standard, and inferring any would be speculation.
Headline overstates, body self-corrects
The quantitative claims are accurate to what is quoted, so the gap is modest rather than large. It is positive because the framing runs ahead of the evidence twice: the title implies that better analyses cause worse follow-through, which the essay concedes nobody has measured, and the prescriptive conclusion that 'only the tracker gets a vote' is asserted without any measurement that tracked engineering remedies reduce recurrence. The transfer of hospital adverse-event findings to software incident reviews is likewise analogical. Explicit self-labeling of the hypothesis and synthesis keeps the overstatement bounded.
Attention incentive, no disclosed commercial stake
The only visible incentive is the ordinary one for a self-published developer-platform essay: a contrarian headline that rewards engagement, and no editorial layer between claim and publication. The supplied text discloses no vendor, product, funding relationship or customer that the argument would sell, and the author repeatedly flags his own synthesis and hypothesis as unproven, which cuts against a purely promotional read. Score reflects venue and framing incentives only; no ownership or funding facts are supplied.
Moderate: solid quotes, one publisher
Confidence is anchored by verifiable named-study quotations and the essay's own explicit epistemic labeling, but limited by having a single publisher, no primary-document access, no adoption signal, and a prescriptive conclusion that the cluster cannot test. The factual claims about the Kellogg study can be assessed with reasonable confidence; the essay's generalization to software cannot.
build
Three manual interventions in a month, and every guard was working as designed1 distinct publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 distinct publisher
build
Force the tool call, then hand Lightsail a long-lived key1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 26, 2026