Invest1 publisher2 min readPublished
MIT and Sakana AI's SIFT finishes a coding agent's self-improvement search on about $34 of API calls
MIT and Sakana AI's SIFT ran a coding agent's full self-improvement search on roughly $34 of API calls, about a tenth of the Darwin Godel Machine's resources. The saving comes from a language-model judge screening patches, so it holds only while that judge picks correctly.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- That Polyglot score came from an o3-mini agent, while the cheap search ran on the open-weight Qwen3-Coder-30B and took 224 CPU hours.
- A regularized Bradley-Terry model turns the judge's head-to-head verdicts into a ranking, and only shortlisted patches go on to the costly full evaluation.
- SWE-60 scores rose from 40.0% to 52.1% over the starting baseline, with a smaller 7.5-point gain on TerminalBench 2.1.
- Xinghong Fu of MIT wrote the paper with Aravinth Kulanthaivelu and Yutaro Yamada, with Sakana AI as partner lab, and posted it on arXiv around September 18, 2026.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- capability A full self-improvement search on an open-weight model comes within reach of teams without deep compute budgets, as Crypto Briefing argues.
- exposure Judge errors can go unnoticed for longer, since the full benchmark only ever scores patches the judge has already shortlisted.
- decision Anyone pricing the method has to test the Qwen3-Coder-30B setup on Polyglot before treating the lead and the low bill as one result.
Thirty steps against 80 is 62.5% fewer [2], and the 4.4-point Polyglot margin [1] is a 14% relative gain on the Darwin Godel Machine's 30.7% [3]. Fewer steps is the weaker half of that claim, or rather, it counts only if a step costs what a node costs. Crypto Briefing's account of the arXiv paper calls SIFT's unit an expansion step and DGM's a node [2]. The two need not be the same size.
The cost line is easier to follow. If the one-tenth figure applies to both compute and API spend, DGM's search ran near 2,240 CPU hours and $340 [4]. SIFT, short for Self-Improvement via Fast Tree-search [1], gets its saving by skipping work. Testing every patch against a full benchmark is the bill that compounds as an agent proposes more candidates [10], and SIFT pays it only for the shortlist its judge produces [14]. Sakana, based in Tokyo, has emphasized evolutionary discovery methods over brute-force strategies [13]. This paper spends its budget on choosing which patches to test [4].
If the judge's picks track what the full benchmark would have said, a search costs about a tenth of what it did [6]. Should the judge drift toward patches that look better and test worse, the search goes wrong at a tenth of the price [12]. A Polyglot margin that owes more to o3-mini than to the search method [3] would make the 4.4 points [1] a result about the model. The account does not say which model produced DGM's 30.7%.
I think the cost claim holds up better than the score claim. The roughly $34 is a whole-search total for one named open-weight setup [6]. The gain on SWE-60 is 30% of its starting score [5], against the 14% Polyglot lead [3]. The counter-case is that beating an agent's own starting point is easier than beating a rival's published number [8], so those two percentages measure different things.
The view fails if the judge's pairwise verdicts often disagree with full benchmark results on the same patches. Matching DGM's quality would then take more full evaluations, and the tenth would not hold [12].
What to watch
- Polyglot scores for the Qwen3-Coder-30B configuration, showing what the roughly $34 search buys on the benchmark where SIFT claims its lead.
- A head-to-head of SIFT and the Darwin Godel Machine on one base model and one budget, separating the method's share of the 4.4-point Polyglot gap from o3-mini's.