Build1 distinct publisher3 min readUpdated
Any write permission useful enough to help an incident is powerful enough to need unwinding, two incident-analysis veterans told an AMA. That reframes the AI SRE buying question.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The asymmetry is what turns read-only from a preference into arithmetic. Long's objection is not that an agent will be wrong more often than it is right. It is that the two outcomes do not weigh the same: she pointed at J. Paul's talk at the same event, where the argument was that bad AI predictions degrade performance far more drastically than good ones improve it [5]. Feed that through a write-enabled tool and a pilot's accuracy rate stops being the interesting number. A correct automated action saves some minutes. An incorrect one plants a second problem inside the first, and the responders now owe time to establishing what changed and whether it can be put back [4].
Speed is the aggravating factor rather than the feature. The unwind bill scales with how many actions got taken before a human noticed, and machine speed is precisely the property being sold. So the vendor's strongest selling point is also the multiplier on its worst case, which is why Long's caveat is not "unless the scenario is low risk" but rather that useful write access and dangerous write access are the same access [3].
Her recommended test is cheap and blunt: run a pilot and ask the engineers, who are overloaded enough to answer honestly and fast [12]. Allspaw, answering a different question about time saved in incident review, put a hole in the obvious version of that test. He cited METR's 2025 developer productivity study: before starting, developers expected AI to make them 24 percent faster; afterwards they believed they had been 20 percent faster; measurement said 19 percent slower [13]. That is a 39 point gap between belief and stopwatch [1], and 43 points between the forecast and the outcome [2]. The study was of developer work, not incident response, and Allspaw offered it as a lens on perception rather than a finding about incident review [16].
Read the two answers together and the pilot survives, with its question narrowed. Ask engineers whether the tool is annoying, whether it interrupts, whether it earns its place in everyday work [11]. Do not ask them how many minutes it saved, because that is the exact quantity the METR numbers say people cannot estimate [13].
Read-only is not neutral either, and Allspaw's questions are the sharp end of that. When a tool sifts data, what does it discard as unimportant? All timelines are opinionated, built from raw data in ways that suited whoever built them, so which events does the machine drop [14]? A summarizer with no production credentials still shapes what responders think they are looking at.
Which leaves the requests that need no credentials at all: helping an incident commander work out who the right SME is, who to page, who has been on the call long enough to be sent away for a break [8]. Long is candid that if the ask came from a board or a VP wanting to be seen using AI, benefit arguments will not move it and you may just have to pick something and let it play out [9][10]. If that is the position, the negotiable surface is not the vendor. It is the permission set, and the tools are in flux enough that nobody can tell you what the right one is yet [15].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
During Incident Fest 2026, John Allspaw and Beth Adele Long of Adaptive Capacity Labs answered audience questions about the relationship between AI and humans in incident response in a virtual AMA hosted by Uptime Labs.
Beth Adele Long said she is a proponent of read-only access for AI during incidents.
Long said any write access that is powerful enough to be useful is also likely to be dangerous, so she declined to carve out an exception for trivial or low-risk scenarios.
Long said incidents are already confusing enough without having to unwind a bizarre decision that was implemented at AI speed.
Long referenced J. Paul's talk at the same event, whose point was that when AI predictions are bad they degrade performance much more drastically than good predictions improve it.
Long said she can see AI helping in long-running incidents as a thought partner, helping responders figure out what is going on and explain current behaviour, in the same way it is already used during routine work.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Verbatim expert opinion, thin empirics
The claims are first-hand and precisely attributable: named practitioners answering named questions in a transcript the publisher itself hosted, so what was said is well established. What is not established is any of it as measured fact — the read-only line is a stated preference, the bad-prediction asymmetry is relayed from another talk with no data, and the single quantitative anchor (METR 2025) is quoted second-hand without link, sample or method, and concerns developer productivity rather than incident work.
No adoption data in sources
The supplied source contains no release, deployment, pricing, licence or usage disclosure — no named AI SRE product, no count of teams granting or withholding write access, and no pilot results. Adoption cannot be measured without inferring facts the source does not provide.
Framing slightly harder than the source
Mildly overstated rather than inflated. The source is one practitioner saying 'I’m a proponent of read-only access,' hedged with the use cases she would welcome; the cluster framing renders that as a drawn line and an answer to the buying question. Allspaw's METR figures are also carried across from developer productivity into incident review with the domain difference noted only in passing. Offsetting this, the practitioners themselves consistently undersell — they flag category immaturity and pose open questions instead of conclusions.
Vendor channel, consultancy guests
Every incentive in the cluster points one way. The publisher sells incident-response capability building and opens the article with a call to action for its own platform; the event is its own. The featured practitioners run a consultancy whose practice is human-centred incident analysis, so a position that keeps humans in the acting role and AI in an advisory one is commercially congruent. No AI SRE vendor, and no team operating write-enabled agents, is given a voice.
Attributable but uncorroborated and unmeasured
Confidence is capped by structure, not by sloppiness: a single interested publisher, zero adoption evidence, one second-hand statistic, and normative advice that cannot be verified from the material. What can be trusted with high confidence is who said what; anything about whether read-only is the right boundary in practice, or whether AI SRE tools save incident-review time, remains untested here.
product
OpenAI prices its own guardrails: 20% more compute, plus a two-week training pause1 distinct publisher
build
The best grade for controlling in-house AI agents is a C+, and buyers can now cite it2 distinct publishers
invest
65,000 pulls a day, one author: the AI coding stack's unpriced dependency1 distinct publisher
science
NIST says AI benchmarks are now an attack surface, not just a measuring stick1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 23, 2026