Build1 publisher2 min readPublished
Google's RRSI gives up training points to keep self-rewritten agent harnesses general
Google Cloud AI Research's RRSI lifted agent scores up to 4.7 points on five unseen benchmarks by capping how far a harness can rewrite itself. Its guardrails cost points on the tuning tasks, a trade worth making for teams that need harness gains to hold on new work.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- According to the paper, much of the recent progress in AI agents comes from work on the harness, the prompts, tools, memory and workflow around a fixed model, and not from new models.
- Harness tuning used to be done by hand, and newer methods have a language model rewrite the harness repeatedly from test-task feedback, a loop the researchers call recursive self-improvement.
- Because that loop keeps working on the same limited tasks, the agent memorizes them, so training scores rise while gains on unseen tasks shrink or disappear.
- The researchers kept the underlying model, Claude Opus 4.8, frozen and compared RRSI with the unmodified baseline harness and four recent optimization methods.
- Every method did well on the training tasks, but on new tasks two of the competing methods ended up below the unmodified baseline harness.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams automating harness tuning need benchmarks the optimizer never sees, since scores on the tuning tasks made every method in this study look like a success.
- cost The token saving is measured against the unregularized optimizer. The untouched baseline harness is leaner still, so adopting RRSI raises per-run spend over not tuning at all.
- exposure Benchmark fitting that never names a task or embeds a solution can pass the critic, leaving the edit budget, the cost rule and held-out scores to catch it.
The paper lists three ways the unregularized search goes wrong. It memorizes patterns that fit only one benchmark, it keeps candidates that scored well by chance, and it piles on complexity that lifts the test score without making the agent better [6]. RRSI puts its controls on the loop and leaves every part of the harness editable [7].
When the loop proposes changes, a budget caps how many independent edits one candidate can bundle [7]. The cap starts wide and narrows. Early rounds allow larger rewrites, and late rounds accept only small changes that can be clearly traced to a result [8]. I think this is the best idea in the method. If a candidate carries six edits and the score moves, there is no way to say which edit moved it, and a lucky combination looks the same as a real fix. The loop also records earlier attempts so it stops chasing ideas that already failed. When progress stalls, it tries parts of the harness it has not yet touched [9].
When it comes to accepting changes, a critic reviews every proposal and rejects any that hardcode task names, solutions or other benchmark-specific tricks [10]. Someone had to write a rule against putting the answer key into the harness. A second rule admits higher compute cost only when a measurable performance gain comes with it, and components that stop helping are removed [11][12]. The regularized harness uses about 30 percent fewer runtime tokens than the unregularized version [16].
The headline scores are maxima. RRSI gained up to 14.1 points on its training tasks [13]. Its best held-out gain, 4.7 points on JobBench, is about a third of that [15][1]. The study used eight benchmarks and held out five, so the optimizer tuned against three [14][2]. On the held-out five, RRSI raised scores in all three domains and never fell below the baseline harness on any of them [21][18].
For those numbers to transfer to another team's agent, three things have to be true. The agent has to run the same frozen model, because the study varied only the harness [22]. The work has to resemble coding, agentic office work or engineering design, the domains the eight benchmarks cover [14]. And the gaps have to be larger than the noise between runs. The report gives the gains as upper bounds and does not include run counts or variance, so the third condition cannot be checked from it [13].
What to watch
- Results for an RRSI-tuned harness moved onto a model other than Claude Opus 4.8, to see whether the gains stay with the harness.
- A release of the RRSI code or critic rules that lets outside teams rerun it against their own held-out tasks.
- Independent reruns on benchmarks outside coding, agentic office work and engineering design.