Build1 distinct publisher3 min readPublished
Published flake numbers from Google, Microsoft, Atlassian and Slack describe an incentive system rather than a discipline failure, and the only lever that changes the arithmetic is a quarantine lane with a named owner.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Start with the arithmetic a developer runs without writing it down. At Google's reported ratio, 84% of transitions from passing to failing involve a flaky test rather than a genuine regression [2]. The Microsoft figure for one investigation is 30 minutes [4]. So the expected value of investigating an unexplained red is 25.2 minutes of confirmed waste against 4.8 minutes of useful work [13]. Retry costs a minute of wall clock and no attention. Hitting it is the correct local decision, which is exactly the dev.to piece's argument: the developer's own experience says 85% of unexplained failures are environment, not code [9].
Slack's mobile suite is where the signal inverts outright. At a 56.76% failure rate [6], a red build is roughly 1.31 times more likely to be noise than a defect [15]. A check that is wrong more often than right is a coin toss with a CI bill attached.
The 150,000 hours a year Atlassian estimated before building detection tooling [5] is the number that will get quoted. Divide by Microsoft's half hour and you get 300,000 investigations a year, about 822 per calendar day [14]. Treat that as a claim about someone else's suite. For it to transfer you would need a comparable count of test executions reaching humans rather than auto-reruns, a similar ratio of flaky to genuine failures, and the same accounting of a developer hour, including the interrupt and not only the triage. The 30 minutes is Microsoft's number and the hours are Atlassian's, so the multiplication is illustrative, not measured.
What follows from this is a design problem, not a values problem. The mechanism I would build: rerun on the same commit to label a test deterministically, move labelled tests to a non-blocking lane, attach a named owner and an expiry date to the label, and let the blocking suite keep failing only on things worth stopping for. Quarantine does not fix flakes. It preserves the meaning of red while someone works through the four root causes the source lists, all of which it calls fixable: async timing, test-order dependency, shared mutable state, and coupling to external services [8]. They stay unfixed because they arrive in no sprint with an owner or a deadline, while user-facing features do [10].
Scope matters, and the source is honest about it. The rational-inaction framing is strongest at 20 or more engineers with separation between developer, QA, infrastructure and management, and it behaves differently under 10 where one person holds several of those seats [11]. If you are a team of six, this is ordinary prioritisation wearing a costume.
The threshold is the useful part of the argument. Once a team accepts that some failures are noise, the bar for investigating drifts upward, and the source puts mild skepticism as early as 5% flakiness [12]. Google's 16% is well past that [1], Microsoft's 25% of failures further still [3]. All four figures come from the same dev.to post, which does not carry primary links in the material supplied [18], so the individual percentages are worth verifying before you put them in a slide. The shape they describe is testable in your own CI in an afternoon. If you cannot state your suite's flake rate to one decimal place, the conversation about engineering culture is premature.
Ranked by verification strength, evidence, and original report placement.
The source describes each role responding correctly on local information: QA logs the failure as likely flaky because without detection tooling it cannot be distinguished from a regression, infrastructure declines scope because flakiness usually is a test quality issue, and managers defer because flakiness does not appear in the product backlog with an owner or a deadline while user-facing features have higher business visibility.
The source characterises the situation as the only common engineering quality problem where every person in the chain responds correctly on local information and the outcome is still catastrophic, and as the final form of Test Debt.
The source states that normalising retries makes a load-bearing decision that test failures no longer reliably indicate real problems, that the threshold for investigating drifts upward once some failures are accepted as noise, and that a suite with 5% flakiness earns mild skepticism; the sentence is truncated in the supplied material.
A flaky test is defined as one that fails intermittently without any change to the code it covers, sometimes passing and sometimes failing with no consistent pattern.
The most common root causes of flakiness are timing issues in async operations, test-order dependencies, shared mutable state, and coupling to external services, and all of these are fixable.
The source states the rational-inaction framing applies most clearly to mid-to-large organisations with role specialisation, is strongest for teams of 20 or more with clear separation between developer, QA, infrastructure and management, and that teams under 10 engineers have different incentive dynamics because one person occupies multiple positions.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single unlinked publisher behind every number
The cluster is one dev.to post. Its qualitative thesis is internally coherent and self-attributed, which is checkable. Every quantitative load-bearing figure is a bare assertion about a third-party organisation with no citation, date, methodology, or link, and one figure is internally inconsistent between two passages. That combination supports the argument's structure but not its arithmetic.
Three secondhand practice disclosures, no outcomes
There are named-organisation practice reports -- Microsoft's owner assignment and merge blocking, Atlassian's detection plus quarantine, Slack's intervention -- which is more than zero adoption signal for the quarantine-and-named-owner pattern. But all three are relayed by the same non-primary publisher, none carries a date, scope, or post-intervention measurement, and no smaller or independent adopter appears. Low, not absent.
Strong quantitative framing on unverifiable inputs
The headline claim that retry wins on expected value past roughly one-in-six flakiness, and the ledger's expected-value and investigations-per-day arithmetic, present decimal-level precision built by multiplying one company's ratio against another company's unit cost -- arithmetic the source itself never performs. Rhetoric runs ahead too: 'only common engineering quality problem', 'catastrophic', 'final form of Test Debt'. The underlying diagnosis is plausible and modestly scoped to 20-plus-engineer organisations, so this is overstatement of precision rather than fabrication.
QA-brand content series with an aligned prescription
The sole source is an instalment in a self-referential numbered series published on a QA-branded dev.to account, and it repeatedly points back to earlier articles about the economic case for test automation. The conclusion it drives toward -- buy or build automated flake detection, quarantine tooling, and ownership enforcement -- aligns with the commercial interest of a QA tooling and consulting brand, and no disclosure accompanies the third-party statistics used to size the pain. This is a visible directional incentive, not evidence of bad faith, and the supplied material contains no funding, sponsorship, or vendor-relationship detail.
Coherent argument, uncorroborated facts
Confidence is limited by structure, not by internal contradiction: one publisher, zero corroboration, no primary links, a truncated remediation section, and one figure used two different ways. The qualitative incentive analysis is stated clearly enough to assess and is honestly scoped to larger organisations, which supports moderate confidence in the framing. The numbers should be re-sourced before any of them are quoted or used in planning arithmetic.
Follow any of these and your For You feed starts watching them — no settings page required.
build
The best review comment asks whether the code belongs, and Microsoft's numbers back it1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
invest
Nvidia's August 26 print: 92% of the quarter rides on one segment1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 27, 2026