Build1 publisher3 min readPublished
Anthropic's hillclimbing loop reverts any agent change that fails on held-out tasks
Anthropic says its claude-api eval workflow, run by Claude Code, lifted the company's own API skill from about 66% to 88%. That gain is self-reported and holds for another team only if its eval is as realistic and as stable as the loop demands.
The Engineer · Build desk

What happened
- Claude Code first builds an eval meant to resemble the work the agent does in production, then runs the optimization against that eval.
- The loop runs in a fixed order: real tasks, a reliable grader, a baseline, one change, a test, and a decision to keep or revert.
- When the strongest setup already scores around 95% or higher, the tooling warns and suggests optimizing cost or latency instead.
- In a customer-support eval, Anthropic's final configuration was more accurate while costing roughly one-fifth as much per ticket.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Adoption costs compute before it saves any, because measuring noise first means repeated baseline runs of the full eval before the first change is scored.
- constraint Teams whose benchmark is a pile of their model's worst failures cannot hillclimb against it meaningfully and have to rebuild it from production-shaped tasks first.
- precedent If the CI framing takes hold, prompt and tool edits to agents get the same gate as code commits, with a suite run on each edit deciding whether it stays.
Hillclimbing an agent with a model has one obvious way to go wrong. The optimizer reads the failures and edits the configuration. The score rises because the edit fits the examples it read. Anthropic's design answers with a split: a change counts only if it holds on examples Claude was never allowed to study [1]. It is the held-out test set from ordinary model training, applied to agent configuration. The workflow runs as two commands, /claude-api build-eval and /claude-api hillclimb [2], and follows a guide written by Anthropic's Lance Martin [10].
Claude measures the eval's noise before it tries to improve anything, according to The Neuron's walkthrough of the guide [7]. The guide's example of an eval too unstable to trust is one where the same system swings between 65% and 90% from run to run [5]. That spread is 25 points [1]. It is a hypothetical in the guide, not a measurement from either of Anthropic's examples. The API-skill gain was 22 points [2]. The explainer does not report the variance measured on the skill eval, so the 22 points mean what Anthropic says only if that eval's spread was well under 22.
A second check tests the eval before anything is optimized against it. Stronger models and more reasoning should generally score higher, and if an expensive frontier model does worse than a smaller one, the task, grader, configuration or environment may be wrong [9]. I think this is the most useful default in the guide. A grader that ranks a small model above a frontier model will send an automated loop after the wrong target.
The guide also warns against mining the benchmark from your own worst runs [8]. Its example runs an agent 1,000 times and keeps the 50 ugliest failures [8]. That is 5% of the runs [3]. Anthropic's explanation is what it calls the jagged capability surface of modern models [8]. A suite built that way tests what one model happens to be bad at. The guide asks for tasks shaped like production; for a support agent, that means refunds, damaged products, billing and account access [12].
Read the customer-support result with the same care. Both results are Anthropic's own examples [4] [14]. That is normal for a methods guide, and it is no evidence about anyone else's ticket queue. For the cost result to carry over, another team's tasks have to resemble what its production traffic actually sends. Its grader also has to give the same answer twice.
The Neuron's explainer flags two limits of its own: automated hillclimbing does not solve the hardest problem, and a better benchmark is not the same thing as a better business [11].
What to watch
- Whether Anthropic publishes the measured run-to-run variance for the API-skill and customer-support evals next to the headline gains.
- Results from teams outside Anthropic running /claude-api hillclimb against their own production tasks and graders.
- Whether the tooling enforces the held-out split itself or leaves it to the user to keep the optimizer away from test examples.