Build1 publisher3 min readPublished
Progressive collapse: the civil engineering word missing from your outage postmortem
Sam Newman's InfoQ talk uses the 1968 Ronan Point tower failure to argue that a cascade is a property of the structure, not of the trigger that started it.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- InfoQ published a presentation by Sam Newman titled 'Understanding Progressive Collapse: How To Avoid A Cascading Failure', in which he applies the civil engineering concept of progressive collapse to distributed digital systems.
- Newman says he came across progressive collapse while doing research for his latest book, while looking at the wider space of resilience engineering and lessons from domains that are not computing.
- Ronan Point, a tower block in Canning Town, suffered a partial collapse shortly after it was opened in 1968.
- The collapse happened at quarter to 6:00 in the morning.
- Quarter to six in the morning is 05:45 on a 24-hour clock.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
InfoQ has published a presentation by Sam Newman that imports a term from civil engineering into distributed systems: progressive collapse [1]. It is a more useful frame than the ones most teams reach for after a cascading outage, because it moves the question from what tripped the system to why the trip had anywhere to go.
Newman says he came across the concept while researching his latest book, reading around resilience engineering and looking for lessons in domains that are not computing [2]. His case study is Ronan Point, a tower block in Canning Town that suffered a partial collapse shortly after it opened in 1968 [3]. The collapse happened at quarter to six in the morning [4], which is 05:45 [15]. Four people died, and Newman argues the number would have been higher at almost any other hour, because most of the rooms that went were living rooms and very few residents were up [5].
The trigger is the part worth sitting with. A resident, Ivy Hodge, put a stove on to heat water, and there was a gas explosion caused by a faulty nut where the oven connected to the main gas supply [6]. Newman describes it as a survivable explosion and not a big one in the grand scheme: Hodge was blown across the room, knocked unconscious, and came to in a puddle of water from the saucepan that had been blown off the top [7]. In service terms, that is a single misconfigured connection in one instance, contained, recovered, no data loss.
What turned it into a building failure was the structure. The explosion blew out the outer wall, which happened to be load-bearing [8]. The four floors above then came down in what Newman calls a concertina effect, taking that corner of the block with them [9]. His definition of progressive collapse is exactly this shape: a small failure produces a significant collapse in the wider system, and the initial trigger looks minor in isolation while cascading into something much larger [10].
Read as design rather than misfortune, Ronan Point has two properties operators will recognise. One wall's removal was enough to unsupport everything resting on it [8][9], and nothing stopped the failure at the boundary of the flat where it began [9]. Newman also makes the position argument: the damage might have been greater had the explosion happened further down the building [11]. Same fault, same blast, different amount of load overhead. That is the difference between a bad deploy in a leaf service and the same bad deploy in the thing everything else calls.
The delivery context is familiar too. Those blocks went up quickly, during a 1960s population boom, in a city where much of Canning Town was still rubble after the Blitz [12]. Urgent demand, fast construction, load paths nobody had reason to test.
One caution about the metaphor, which Newman raises himself. He compares the effect to a domino demonstration where a small domino topples a larger one, and notes that in that kind of system you can see all the moving parts and reason about them [14]. Distributed systems are the case where you cannot, which is why the structural question has to be asked before the incident rather than reconstructed after it.
Newman says the talk goes on to two more failures: one he was intimately involved with, and one he suggests may have affected many in his audience last year [13]. Worth checking whether those cases produce a testable rule for load paths and blast boundaries, or stay at the level of a good analogy.