Invest1 publisher2 min readPublished
Google's open-source RRSI rewrites agent scaffolding to add six points on Terminal-Bench
Google Cloud AI Research has released RRSI, an Apache 2.0 tool whose self-rewriting agent harness lifted Terminal-Bench 2.1 scores from 74.2% to 80.2%. Any team can use it commercially, though on tasks the agent never trained against the reported gain falls to 4.7 points.
The Investor · Invest desk

What happened
- RRSI holds the model fixed while the agent edits its own prompts, control flow, tools and memory, tests each change, keeps what works and starts the next round from the improved version.
- Google Cloud AI Research built the framework with UNC-Chapel Hill, Stanford and Washington University in St. Louis and introduced it on September 21, 2026.
- Across eight benchmarks the researchers reported gains of up to 14.1 points on the splits the harness was evolved on.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- cost Adopting RRSI adds a search bill on top of inference. Its 30% policy-token saving is measured against other evolution methods, so it makes searching cheaper without making it free.
- exposure On the researchers' own reported figures, the best tuned-split gain is three times the unseen-task gain, so a buyer shown RRSI-evolved scores on tuned splits could be overpaying for that gap.
- precedent With the code under Apache 2.0, rival agent builders can adopt the same leakage and noise filters, so buyers can reasonably expect any self-improving harness to report scores on unseen tasks.
Counted as failures, the Terminal-Bench 2.1 result the researchers reported is bigger than six points sounds. An agent at 74.2% fails 25.8% of tasks and one at 80.2% fails 19.8% [8], so the evolved harness cleared about 23% of the misses [2]. SWE-bench Verified shows less on the same count. Failures there fell from 18.0% to 16.2%, a tenth of what was left [3].
The more useful comparison is between splits. Set the best case on the evolve splits, 14.1 points [10], against the five out-of-distribution benchmarks the agent did not train against, where Crypto Briefing reports an improvement of 4.7 points [11]. The unseen-task figure is a third of the best tuned one [1]. The two are not like for like: 14.1 is a maximum, and the account does not say whether 4.7 is an average across the five or another best case.
A gap between tuned and unseen tasks is what benchmark-gaming looks like from outside, and two of RRSI's four selection filters target it. The leakage critic screens for edits that carry benchmark-specific knowledge into the harness, and the noise floor ignores gains too small to tell apart from random variation [6]. Proposals are constrained as well. The agent is allowed sweeping edits early, is pushed toward small tweaks as the run matures, and conditions each new proposal on what it has already tried [5]. A 1.8-point move on SWE-bench [9] is the size of gain a reader would want to test against that floor's setting in the paper itself, posted on arXiv as 2609.24972 [14].
All the spending goes on scaffolding. The model's weights stay fixed throughout [1], and harnesses evolved on Gemini 3.5 Flash improved results when paired with smaller Gemini models [13]. A team could run the search once and deploy the result on a smaller model without paying to train anything.
I think the 4.7-point figure is the one to plan around. At that size, RRSI's commercial case rests on whether a search run costs less than the engineer time spent tuning prompts and tools by hand. The counter-case is breadth. The tests covered coding on Terminal-Bench, workspace tasks on Harvey LAB and engineering design on EngDesign [12], which argues against the gains being one benchmark's quirk. I would be wrong if 4.7 points turns out to be an average across all five unseen benchmarks with no single one carrying it. A third of the best tuned gain, on tasks nobody tuned for, would then be a good return on a loop the Apache 2.0 licence lets any team run commercially [3].
What to watch
- Outside reproductions of the five out-of-distribution results using the GitHub code Google published alongside the arXiv paper.
- Tests of RRSI-evolved harnesses on non-Gemini models, beyond the smaller Gemini variants reported so far.