Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Anthropic's on-call engineer opens Claude before the dashboard, and is still hiring humans

Alex Palcuie says he now goes to the model first when Claude pages him, and that the answer to whether Claude fixes its own incidents is no. His evidence is his team's hiring plan.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying Anthropic's on-call engineer opens Claude before the dashboard, and is still hiring humans
Photo: qconlondon.com

What happened

  • An Anthropic reliability engineer says that since about January he has opened Claude before his monitoring dashboards during incidents.
  • Asked directly whether Claude fixes his incidents, his answer is no, and he says claiming otherwise would be hypocritical.
  • His stated evidence is that the team exists at all and is hiring in London, Dublin and the US, including staff-level roles.
  • His complaint about the benchmarks is that they score models on problems already packaged into a clean prompt.

Why it matters

  • contradiction Quote either half of this talk on its own and you get opposite buying advice, because the same engineer trusts the model in the first minutes of a page and is funding humans for everything after it.
  • constraint If solve rates are measured on pre-formatted incidents, procurement has no number that predicts behaviour on a real page, and the hardest part of triage stays unpriced.
  • exposure Any AI SRE vendor now has to explain why its product works where the model maker's own reliability team, with no token budget to worry about, reports that it does not.
  • decision For engineers whose leadership wants an AI ops purchase, this reframes the choice as pointing an existing model at first-response context gathering rather than buying an owner for the pager.

Separate triage from ownership and the two halves of Palcuie's account stop fighting each other. Going to the model before the dashboards [4] is a claim about the first minutes of a page: what changed, and where to look next. Recruiting staff-level engineers across three countries [7] is a claim about who owns the fix and the follow-up. He does not say the first is eating the second, and he is blunt that it would be hypocritical to pretend otherwise [5].

The hiring line is worth more than any benchmark number because it costs money. Anthropic is the most favourable test case anyone could construct for automated incident response: unlimited tokens and researchers a desk away, with the on-call engineer himself involved in training the models [8]. Palcuie sets that up deliberately before answering no [5]. A demo runs on a curated dataset [11]. A headcount plan runs on a forecast someone has to fund, and by his own account the team is barely a year old and still growing [6][7].

His objection to the benchmark industry is the part worth taking away. What gets scored is whether a model can solve a problem that has already been packaged into a clean prompt, and he says that is not what a 3 a.m. page looks like [12]. The packaging is the work. Deciding that the latency graph and the deploy log and a customer complaint are the same event is the expensive judgement, and a scoring harness that hands the model a tidy problem statement has already done it for free.

Meanwhile the supply side keeps growing. Palcuie counts at least ten companies pitching some version of AI SRE [9], with two more surfacing while he wrote the slides [10], which puts him at twelve or more before he had delivered the talk [19]. He angel invests and says the pitches make him more skeptical rather than less [16], while also saying he wants those companies to keep practising [15]. That is a coherent position, and it is the one being asked about at dinner parties by engineers whose VPs have told them to do something with AI [17].

Two things the material does not settle. The transcript we have breaks off just as he turns to the useful ways Claude helps him on call, so the specific tasks he trusts it with are not in evidence here [18]. And his no carries an asterisk about timelines: he says he would not be surprised if this becomes possible later [14]. The load is not theoretical either. He says Claude goes down more often than anyone would like, and he was pulled into an incident while at the conference [13].

What to watch

  • Whether the full talk names the specific triage tasks Palcuie trusts the model with, since the published excerpt stops before that list.
  • Whether Anthropic's reliability hiring slows, given that headcount is the evidence he himself offers for the answer being no.
  • Whether any AI SRE benchmark starts scoring on raw, unpackaged incident signals instead of curated prompts.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence42
Adoption34
Hype gap+28
Incentives66
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Alex Palcuie is on the AI reliability team at Anthropic and describes his job as keeping Claude up.

    ReportedSupportedSource: Alex Palcuie, InfoQ presentation transcriptView cited source
  2. [2]

    Palcuie did on-call for Claude's serving stack, Solo, for his first three months, then onboarded everyone else so he could stop being on-call; he was one of two people who joined the reliability team in London.

    ReportedSupportedView cited source
  3. [3]

    Before Anthropic, Palcuie was an SRE at Google on Google Cloud's compute product GCE, on the 'SRE for SRE' team, the escalation layer for outages bad enough that normal SREs want backup.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. infoq.com

    1 article · August 26, 2026

    Presentation: Can Claude Fix Itself? Using LLMs for Incident Response

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories