Skip to content

Build1 publisher3 min readPublished

The incident review timeline assumes one fix that holds

Vanessa Huerta Granda, who leads resiliency engineering at Enova, spends her InfoQ talk on incidents where the fix gets announced and then withdrawn. The controls she names are rotation, breaks and room capacity.

The Engineer · Build desk

Photograph accompanying The incident review timeline assumes one fix that holds
Photo: infoq.com

What happened

  • In an InfoQ presentation titled When Incidents Refuse to End, Enova resiliency engineering lead Vanessa Huerta Granda says the trained timeline of break, fix, back to normal does not describe real incidents.
  • She puts the really long incidents in a category of their own, the marathons, and says they show the difference between work as imagined and work as done.
  • Her list of what actually gets asked mid-incident includes who owns a component, a question she calls a good one precisely when nobody knows the answer.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A three-state review timeline has one slot for the fix, so the stretch of a long incident spent believing a wrong resolution is the part the retrospective is worst at recording.
  • decision Once responder cognition sets the ceiling, somebody has to own a rotation plan before the incident starts, and the plan has to say what state moves with each swap.
  • exposure An unowned component becomes reachable as a delay: the ownership lookup happens on the clock, with the system already degraded, and no telemetry answers it.

A review timeline with one fix step has one slot for the fix. Huerta Granda's account of a real incident needs at least two: the fix that gets announced, and the fix that holds. "Wait, we fixed it. Wait, no, we didn't actually fix it," she said of how the middle of an incident goes [5]. The hours between those two sentences are real response work, and a diagram that runs break, fix, normal has nowhere to put them [2].

Ownership is the other thing the template assumes is already settled. She listed the question of who owns a component among the ones that come up mid-incident, and said, "That's a really good one when you don't know who actually owns that" [4]. That answer is not in the telemetry. It sits in a service catalog, a git history, or somebody's memory, and the lookup runs while the system is still degraded.

The controls she names for the long ones are not diagnostic at all. Her incident commander duties have included reminding people to take bathroom breaks and to "rotate responders before they lose brain function", and opening the doors of a war room that was built for 10 people [12][13]. She also joked that she liked incidents that ran through lunchtime, because the meal became expensable [14]. InfoQ's published transcript breaks off mid-sentence in the war room passage, before any handoff or review practice [13].

Rotation is a handoff. Each swap moves the state of an investigation, including every hypothesis already ruled out, from one head to another, and a long incident has several of them. She said these incidents "show us the difference between work as imagined versus work as done" [7].

Treat the category as one practitioner's population. Huerta Granda leads the resiliency engineering team at Enova, was the sole incident commander on call for four years in her 20s, and describes herself as an industrial engineer by training who has not coded in years [8][9][11]. The talk is process observation from inside incident command, not a measured sample of incidents. For the marathon category to change anything where you work, your incidents have to run long enough that fatigue and rotation affect the outcome. If the median one closes inside an hour, rotation is not your limit, though the ownership lookup can still cost you wall-clock time.

The change she dates is in what organizations ask for. Over the past five to eight years, she said, they stopped wanting only quick fixes: "They want lasting resilience" [15]. That window sits inside the decade she says she has spent in incident response, and part of it she spent away from Enova working with teams across the industry [16][10].

What to watch

  • Whether InfoQ publishes the rest of the transcript, including the war room passage that breaks off and whatever handoff practice follows it.
  • Whether the Resilience in Software Foundation puts incident duration data behind the marathon category instead of practitioner accounts.
  • Whether incident tooling adds a state for a fix that was announced and then withdrawn, so retrospectives can record it.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories