Build1 publisher3 min readPublished
NVIDIA let an optimizer agent find the four token savers in its SoL-Pi harness for Pi
NVIDIA released SoL-Pi, an MIT-licensed Pi extension with four token-saving harness mechanisms that an optimizer agent found. A dev.to review puts the saving near one third on long sessions, bought with a few lost solves.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- An optimizer agent read execution traces from a base Pi harness, proposed changes, implemented them and validated each one in real environments.
- The search covered about 150 directions in six proposal families, roughly 535 environments, more than 3,000 runs and over 60,000 agent-environment interactions.
- Capability tolerances were fixed before the search and hidden from the optimizer, and the EdgeBench set was frozen and held out so its results never fed back.
- In a swarm test reported by the review, 20 SoL-Pi workers reached better optimization results than 20 Pi workers for 26.8% less cost.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- contradiction The review's summary of about 94% of quality sits beside a Terminal-Bench 4 result of 83% of baseline solves, so the quality cost depends on which benchmark resembles a team's own work.
- exposure Turning on the reducer sends eligible build and test logs to whatever model is configured, and the project's SECURITY.md says not to enable it on logs that must stay on the machine.
- constraint Short sessions and local or free models gain nothing, because the mechanisms trigger only when context accumulates and the savings are counted in API spend.
- capability SoL-Pi adds its mechanisms through Pi's public APIs without patching Pi, so teams have a working template for writing their own harness changes.
Two of the four mechanisms add no model calls [14]. Action fusion lets an edit or write call carry a then_run step, a test or a build, inside the same call [7]. That removes the extra model round trip the agent would otherwise spend asking for the run [7]. ObservationPack archives any tool result over 10 KiB on local disk. It sends the result in full for two requests, then replaces it with a stable handle and a head-and-tail excerpt [9].
The compaction policy is the one I would copy into another harness. At each plan-step completion it estimates the requests still to come against the cost of rewriting the prompt cache, and compacts only when the projected savings win [8]. Near the context window limit it compacts anyway [8]. The policy treats compaction as a cache rewrite with a price, so a session with few requests left keeps its cached prefix [8].
The reducer puts a second model in the loop, and it is guarded to match. Build and test logs of 4 KiB or more go to a cheap model, GPT-5.6 Luna, which extracts a receipt [10]. A deterministic verifier then checks schema, hash, exit status, exact quotes and size, and any failure sends the original log instead [10]. This is careful work. The cheap model's summary reaches the agent only after a check that cannot hallucinate has passed it [10]. The reducer runs ahead of ObservationPack, and ObservationPack recognizes its marker so verified evidence is not processed twice. The originals stay on disk [11].
The search ran as a broad-to-deep funnel [5]. An outer loop kept many isolated, disposable lineages that were cheap to kill. An inner loop cycled through implementation, independent review and revision [5]. The title says "Recursively". According to the review, the paper itself calls "recursive efficient improvement", where a cheaper harness makes the next harness search cheaper, a long-term vision it has not demonstrated [17].
The one-third figure comes from the case the mechanisms were built for: a 200-turn session where re-read context dominates the bill, according to the review [12]. Terminal-Bench 4 is a harder test of the trade. If cost per solved task is total spend divided by solves, SoL-Pi spent about $211 and stock Pi about $286 [1]. SoL-Pi's run cost about 26% less [3]. Stock Pi's three extra solves cost about $25 each at the margin, well above either harness's average per solve [4]. I'd keep the stock harness for any work where the next solved task is worth more than $25 [4].
Teams that do not want logs going to a second model can start with the two mechanisms that add no calls. The review says that conservative configuration, actionFusion plus observationPack, gets much of the saving with no extra model calls and no run interruption [14]. It does not put a number on "much".
What to watch
- Independent runs of SoL-Pi on other benchmarks or team workloads that test whether the one-third saving holds outside re-read-heavy 200-turn sessions.
- Reported fallback rates for the reducer's verifier, since a high rate would mean most logs reach the agent at full size anyway.
- Any demonstration of the 'recursive efficient improvement' loop, where a cheaper harness measurably pays for the next harness search.