Build1 distinct publisher3 min readPublished
Alex Palcuie says he now goes to the model first when Claude pages him, and that the answer to whether Claude fixes its own incidents is no. His evidence is his team's hiring plan.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Separate triage from ownership and the two halves of Palcuie's account stop fighting each other. Going to the model before the dashboards [3] is a claim about the first minutes of a page: what changed, and where to look next. Recruiting staff-level engineers across three countries [6] is a claim about who owns the fix and the follow-up. He does not say the first is eating the second, and he is blunt that it would be hypocritical to pretend otherwise [4].
The hiring line is worth more than any benchmark number because it costs money. Anthropic is the most favourable test case anyone could construct for automated incident response: unlimited tokens and researchers a desk away, with the on-call engineer himself involved in training the models [7]. Palcuie sets that up deliberately before answering no [4]. A demo runs on a curated dataset [10]. A headcount plan runs on a forecast someone has to fund, and by his own account the team is barely a year old and still growing [5][6].
His objection to the benchmark industry is the part worth taking away. What gets scored is whether a model can solve a problem that has already been packaged into a clean prompt, and he says that is not what a 3 a.m. page looks like [11]. The packaging is the work. Deciding that the latency graph and the deploy log and a customer complaint are the same event is the expensive judgement, and a scoring harness that hands the model a tidy problem statement has already done it for free.
Meanwhile the supply side keeps growing. Palcuie counts at least ten companies pitching some version of AI SRE [8], with two more surfacing while he wrote the slides [9], which puts him at twelve or more before he had delivered the talk [12]. He angel invests and says the pitches make him more skeptical rather than less [16], while also saying he wants those companies to keep practising [15]. That is a coherent position, and it is the one being asked about at dinner parties by engineers whose VPs have told them to do something with AI [17].
Two things the material does not settle. The transcript we have breaks off just as he turns to the useful ways Claude helps him on call, so the specific tasks he trusts it with are not in evidence here [18]. And his no carries an asterisk about timelines: he says he would not be surprised if this becomes possible later [14]. The load is not theoretical either. He says Claude goes down more often than anyone would like, and he was pulled into an incident while at the conference [13].
Ranked by verification strength, evidence, and original report placement.
Alex Palcuie is on the AI reliability team at Anthropic and describes his job as keeping Claude up.
Palcuie did on-call for Claude's serving stack, Solo, for his first three months, then onboarded everyone else so he could stop being on-call; he was one of two people who joined the reliability team in London.
Before Anthropic, Palcuie was an SRE at Google on Google Cloud's compute product GCE, on the 'SRE for SRE' team, the escalation layer for outages bad enough that normal SREs want backup.
Palcuie says that since about January this year he has started reaching out for Claude before he reaches out to his monitoring dashboards, something he calls slightly transgressive to admit, having been naturally skeptical at the beginning.
Asked whether Claude really fixes his incidents, Palcuie's answer is no, and he says it would be genuinely hypocritical for him to stand up and say Claude fixes everything.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One first-person practitioner transcript, no metrics
All claims trace to a single partial conference transcript from one publisher. The speaker is unusually well positioned - he ran on-call for Claude's serving stack and previously worked Google's SRE-for-SRE escalation layer - and his central negative claim is against his employer's commercial interest, which raises credibility. But there are no uptime, MTTR or benchmark numbers, no named vendors or benchmarks, no corroborating source, and the excerpt breaks off before the promised account of what the model does well.
Disclosed internal assist, no autonomous operation
There is real disclosed usage - a model-first triage habit inside the vendor's own reliability team since roughly January, plus an at-least-12-company vendor field with benchmarks and datasets. But adoption stops well short of the headline capability: no autonomous incident resolution, no paging handed to a model, and continued hiring of human on-call staff in three geographies. No counts, seat numbers or customer deployments are given for any AI SRE product.
Category claims run ahead of the practitioner record
The article itself is deflationary and roughly aligned with its evidence, so the gap sits in the surrounding AI SRE narrative it describes: at least a dozen companies, benchmarks, datasets and papers scoring percentage-of-incidents-solved, against an insider account that the model does not fix incidents even where conditions are most favourable and where benchmarks are said to test clean pre-packaged prompts rather than real pages. The gap is positive but moderate, not extreme, because the speaker explicitly qualifies the 'no' as a timeline statement and says he expects the capability eventually.
Two disclosed conflicts, both pointing against the message
The speaker works for the vendor whose model is under discussion, and separately discloses angel investing, both of which would normally bias toward optimism about AI incident response. Instead he states the capability does not work and says he is more skeptical of AI SRE startups than people expect, which is a counter-incentive pattern. The publisher relationship is also structural: this is conference presentation content reproduced as a transcript, so speaker framing carries through with limited editorial challenge, and no vendor named in the category gets to respond.
Credible insider testimony, single-source and incomplete
Confidence is moderate: the speaker is directly accountable for the system he describes and his central claim runs against his employer's interest, which is a strong signal. It is capped by single-publisher sourcing, the absence of any quantitative measure, unnamed vendors and benchmarks, and a transcript that stops mid-sentence before the balancing material.
product
Leaderboards as a procurement trap: when the test rig outranks the model1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
build
Perf work stopped being a specialist queue item, and slow endpoints became a choice1 distinct publisher
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026