BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Anthropic's on-call engineer opens Claude before the dashboard, and is still hiring humans
Alex Palcuie says he now goes to the model first when Claude pages him, and that the answer to whether Claude fixes its own incidents is no. His evidence is his team's hiring plan.
The Engineer · Build desk

What happened
- An Anthropic reliability engineer says that since about January he has opened Claude before his monitoring dashboards during incidents.
- Asked directly whether Claude fixes his incidents, his answer is no, and he says claiming otherwise would be hypocritical.
- His stated evidence is that the team exists at all and is hiring in London, Dublin and the US, including staff-level roles.
- His complaint about the benchmarks is that they score models on problems already packaged into a clean prompt.
Why it matters
- contradiction Quote either half of this talk on its own and you get opposite buying advice, because the same engineer trusts the model in the first minutes of a page and is funding humans for everything after it.
- constraint If solve rates are measured on pre-formatted incidents, procurement has no number that predicts behaviour on a real page, and the hardest part of triage stays unpriced.
- exposure Any AI SRE vendor now has to explain why its product works where the model maker's own reliability team, with no token budget to worry about, reports that it does not.
- decision For engineers whose leadership wants an AI ops purchase, this reframes the choice as pointing an existing model at first-response context gathering rather than buying an owner for the pager.
Separate triage from ownership and the two halves of Palcuie's account stop fighting each other. Going to the model before the dashboards [4] is a claim about the first minutes of a page: what changed, and where to look next. Recruiting staff-level engineers across three countries [7] is a claim about who owns the fix and the follow-up. He does not say the first is eating the second, and he is blunt that it would be hypocritical to pretend otherwise [5].
The hiring line is worth more than any benchmark number because it costs money. Anthropic is the most favourable test case anyone could construct for automated incident response: unlimited tokens and researchers a desk away, with the on-call engineer himself involved in training the models [8]. Palcuie sets that up deliberately before answering no [5]. A demo runs on a curated dataset [11]. A headcount plan runs on a forecast someone has to fund, and by his own account the team is barely a year old and still growing [6][7].
His objection to the benchmark industry is the part worth taking away. What gets scored is whether a model can solve a problem that has already been packaged into a clean prompt, and he says that is not what a 3 a.m. page looks like [12]. The packaging is the work. Deciding that the latency graph and the deploy log and a customer complaint are the same event is the expensive judgement, and a scoring harness that hands the model a tidy problem statement has already done it for free.
Meanwhile the supply side keeps growing. Palcuie counts at least ten companies pitching some version of AI SRE [9], with two more surfacing while he wrote the slides [10], which puts him at twelve or more before he had delivered the talk [19]. He angel invests and says the pitches make him more skeptical rather than less [16], while also saying he wants those companies to keep practising [15]. That is a coherent position, and it is the one being asked about at dinner parties by engineers whose VPs have told them to do something with AI [17].
Two things the material does not settle. The transcript we have breaks off just as he turns to the useful ways Claude helps him on call, so the specific tasks he trusts it with are not in evidence here [18]. And his no carries an asterisk about timelines: he says he would not be surprised if this becomes possible later [14]. The load is not theoretical either. He says Claude goes down more often than anyone would like, and he was pulled into an incident while at the conference [13].
What to watch
- Whether the full talk names the specific triage tasks Palcuie trusts the model with, since the published excerpt stops before that list.
- Whether Anthropic's reliability hiring slows, given that headcount is the evidence he himself offers for the answer being no.
- Whether any AI SRE benchmark starts scoring on raw, unpackaged incident signals instead of curated prompts.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence42
- Adoption34
- Hype gap+28
- Incentives66
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Alex Palcuie is on the AI reliability team at Anthropic and describes his job as keeping Claude up.
- [2]
Palcuie did on-call for Claude's serving stack, Solo, for his first three months, then onboarded everyone else so he could stop being on-call; he was one of two people who joined the reliability team in London.
- [3]
Before Anthropic, Palcuie was an SRE at Google on Google Cloud's compute product GCE, on the 'SRE for SRE' team, the escalation layer for outages bad enough that normal SREs want backup.
- [4]
Palcuie says that since about January this year he has started reaching out for Claude before he reaches out to his monitoring dashboards, something he calls slightly transgressive to admit, having been naturally skeptical at the beginning.
- [5]
Asked whether Claude really fixes his incidents, Palcuie's answer is no, and he says it would be genuinely hypocritical for him to stand up and say Claude fixes everything.
- [7]
Palcuie cites the fact that his team exists and is hiring for many positions in London, in Dublin and in the US, including staff positions, as showing that an LLM cannot carry a pager; he says if an LLM could, they might not need to hire so much.
- [8]
Palcuie notes that Anthropic has unlimited tokens, that researchers sit a desk away from him, and that he gets involved in training the models, framing his team as the most likely place for automated incident response to work.
- [9]
Palcuie says he knows at least 10 companies right now whose entire pitch is some version of AI SRE.
- [10]
Palcuie says at least two other AI SRE companies appeared by the time he was writing his slides, and that his list is not exhaustive.
- [11]
Palcuie says the AI SRE space now has benchmarks, curated datasets, historical incidents, papers, and a game of what percentage of those incidents an app can solve.
- [12]
Palcuie says some benchmarks amount to asking whether the model can solve problems pre-packaged into a clean prompt, which is not what something looks like when you get paged at 3 a.m., because actual incidents are not well-formatted problems.
- [13]
Palcuie says Claude is down more often than any of them would like, and that he was involved in an incident earlier even though he was at the conference; people tag him on social media when Claude goes down.
- [14]
Palcuie puts an asterisk on his no, saying it relates to timelines and that many people would not be surprised if at some point in the future this became possible.
- [15]
Palcuie says he is not cynical about the goal, that he is genuinely cheering for the AI SRE companies, and that he wants them to practice.
- [16]
Palcuie says he is open about angel investing and that people expect him to be less skeptical about AI SRE startups, but he is more skeptical.
- [17]
Palcuie says he is asked whether Claude fixes his incidents at dinner parties, in conference hallways, and by friends at his old company whose VPs are telling them to do something with AI.
- [18]
The supplied transcript excerpt breaks off mid-sentence before Palcuie enumerates the useful ways Claude helps him during his on-call, which he says earlier in the talk that he will do.
- [19]
By Palcuie's own count there were at least 12 AI SRE companies before he delivered the talk.
Sources
1 independent publisher whose own reporting we read for this story.
- infoq.comPresentation: Can Claude Fix Itself? Using LLMs for Incident Response
1 article · August 26, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.