Skip to content

Build1 publisher2 min readPublished

Picking among eight candidate commands per turn lifts a 9B terminal agent from 50% to 68%

NVIDIA and KAIST researchers lifted a 9B terminal agent from 50.00% to 68.03% Pass@1 by letting a stronger model pick among eight commands per turn. A 9B judge trained on that model's preferences reached 57.14%.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Picking among eight candidate commands per turn lifts a 9B terminal agent from 50% to 68%
Generated illustration

What happened

  • In the writeup's example, a terminal agent runs pip install yaml instead of pip install pyyaml on turn two, then wrecks its own environment trying to recover.
  • Mid-Harness places a verifier between the policy model and the shell, so only the verifier's pick from the sampled candidates is executed and the rest are discarded.
  • Putting all eight candidates in one prompt and asking the 9B model to pick a letter reached 51.02% Pass@1, about a point above baseline.
  • A ring-duel scheme built on four pivots avoids the 28 duels of a full round-robin and reached 54.76% with the undistilled 9B model as judge.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams adopting the method choose between 10.89 more points of Pass@1 and a call to a stronger model on every turn of every session.
  • capability A wrong command is discarded before it can corrupt a lockfile, kill a daemon or half-migrate a schema, damage the writeup says prompting cannot easily undo.
  • cost Each action waits for a verdict before it reaches the shell, so selection latency is added to every turn, while Best-of-N spends its extra compute on parallel sessions and containers.

According to a writeup of the NVIDIA and KAIST paper (arXiv:2609.39982), the central finding is about ranking [1]. Code models already generate a correct bash command on turns where their greedy output is wrong. A small model's weakness is putting that command first [12]. The 9B generator was not fine-tuned, and the harness was not changed [7]. All 18.03 points gained with GPT-5.6 Sol as verifier came from choosing among commands the 9B model had already sampled [1][6][7].

The writeup goes further. It says the paper proves that re-running entire trajectories from scratch is an expensive misallocation of compute [15]. Its own Best-of-N example is six parallel thirty-turn sessions, or 180 policy turns [3][3]. Pairwise selection at eight candidates costs eight samples plus 22 verifier duels per turn [4][10]. That is 30 model calls a turn and 900 over a thirty-turn session, five times the call count of the re-run example [4][6].

Call counts are not spend. Three things would have to hold for per-turn selection to come out cheaper. The eight samples are drawn from one shared history [4], so a serving stack that caches that prefix processes the history once per turn instead of eight times. Duel prompts and verdicts have to stay short. And rejecting a bad command early has to shorten sessions by more turns than verification adds. Re-runs, for their part, pay for a fresh container per attempt [3]. The writeup's results are Pass@1 rates [16]. It does not include token, latency or dollar costs, or a whole-session Best-of-N baseline on TerminalBench-Lite.

The selector results are the part I would copy. Pointwise scoring, eight independent zero-to-ten ratings, reached 52.38% [9]. The writeup attributes the weak listwise result to language models struggling to compare several bash commands presented as one list [17]. Anyone who has reviewed an eight-way diff will sympathise. Pairwise duels gained 4.76 points to pointwise's 2.38, twice the gain for 2.75 times the verifier calls [5][7][8].

I think the narrow claim holds on this evidence. At 9B, picking among sampled commands before they run beat executing the greedy one under every selection method the writeup reports [5][8][9][10][11][6].

What to watch

  • The paper's ablation on whether the verifier must write a chain-of-thought rationale before judging, since rationale length sets the cost of each duel.
  • A matched-compute comparison against whole-session Best-of-N on TerminalBench-Lite, measured in tokens or wall-clock time.
  • Results with policy models larger than TMAX-9B, where the greedy command may already rank first more often.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories