Skip to content

Build1 publisher3 min readPublished

Google's Dream-RSI cuts Gemini calls 42% on a Lasso task by evolving only its search policy

Google researchers' Dream-RSI cut Gemini calls on a Lasso solver task from 550 to 317 by rewriting a Python search policy, with every model frozen. Teams running scored code search can test it as a call-budget saving on their own tasks.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Google's Dream-RSI cuts Gemini calls 42% on a Lasso task by evolving only its search policy
Photo: dream-rsi.com

What happened

  • The evolved policy decides which branches of the search tree to expand, how many run in parallel, how deep to go and when to stop.
  • On the same Lasso task, the solver produced with the evolved policy ran in 2.93 seconds against 3.59 seconds.
  • Dream-RSI's circle-packing score of 2.635983 matched AlphaEvolve V2's 2025 result to six decimal places.
  • The paper comes from 17 authors at Google, Google DeepMind, the University of Maryland and the University of Virginia.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost The saving comes off the coder-call bill, so teams whose code-search spend sits mostly in evaluation will recover a smaller share of their total.
  • constraint Gains come from steering a frozen coder's search, so the ceiling is still what Gemini can propose, and the circle-packing tie with AlphaEvolve V2 fits that limit.
  • decision Adopting Dream-RSI means changing the scheduler on an existing code-search loop, so the case for it rests on a team's own call counts and scorer costs.

Each search runs through one method. In the minimal structure printed in Appendix B.2, `solve()` resets the question and then loops: read what has been revealed so far, update the set of closed branches, select a batch, probe it, record the result curve [7]. An empty batch breaks the loop. The loop condition is `while not _budget_done(question, budget):`, so the policy runs until the budget is spent [7].

The system rewrites only that policy [4]. Gemini, working through Gemini CLI, writes candidate code. A fixed evaluator scores it, and a second frozen LLM rewrites the policy [5]. In the abstract, the authors wrote: "A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying coding agent unchanged." [12]

Google DeepMind announced AlphaEvolve in May 2025. It already proposed programs, scored them, kept the best and mutated them, with the model held fixed [14]. Dream-RSI adds one layer on top of that kind of loop [14].

The headline cut checks out: 233 fewer calls out of 550 is 42.4% [1]. The runtime gain on the same task works out to about 18% [2]. On Lasso, then, the policy got a better program out of less search. The circle-packing tie with AlphaEvolve V2 is harder to price, because the dev.to write-up does not give a call count for that run [10].

Eliezer Yudkowsky fixed the safety community's meaning of the term in a 2008 LessWrong essay. In that sense a seed AI rewrites its own source or weights, and the thing that improves is the thing doing the improving [13]. Here a different model rewrites the policy, and that model never changes. According to the dev.to write-up, the model's weights never move anywhere in the 36-page paper [3][1]. The X posts saying Google had "cracked" recursive self-improvement [2] appear to have stopped at the title. The write-up's own verdict: "A discount; nothing here looks like takeoff." [16]

I think that verdict is right, and the discount is still worth having. The 42% is a claim about one Lasso task scored by a fixed evaluator [8][5]. Three things have to hold before it transfers to another team's loop. The scorer has to be cheap enough to run on every candidate. Coder calls have to dominate the bill. And the search tree needs branches worth closing early.

The design choice worth copying is making exploration a program. A Python policy can be read and diffed. Because the coding agent is left unchanged [12], the policy can sit on an existing loop without retraining anything. The dev.to write-up suggests asking three things of any "RSI" paper: what object changes, who is frozen, and what the headline number is a unit of [15]. For Dream-RSI the answers are a Python search policy, all three models, and Gemini calls [4][5][8].

What to watch

  • A call count for the circle-packing run would show whether matching AlphaEvolve V2 also came cheaper.
  • Teams rerunning the Lasso setup on their own scorers, and whether their call savings land anywhere near 42%.
  • A follow-up in which the coder or policy agent updates its own weights would meet the 2008 definition the safety field uses.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories