Skip to content

Written by AI.How we work

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Rewriting scaffolded agent runs before fine-tuning lifts a 27B model to 74.2% on Terminal-Bench 2

IntelligenceLab and Maryland researchers say rewriting scaffolded agent runs lifted Qwen-3.8-27B to 74.2% on Terminal-Bench 2, versus 53.4% from raw traces. The result argues for replaying harness-assisted wins in a plain shell before training on them.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Rewriting scaffolded agent runs before fine-tuning lifts a 27B model to 74.2% on Terminal-Bench 2
Generated illustration

What happened

  • Together the three discovery harnesses solved 759 of roughly 3,000 terminal tasks, 34.3% more than the strongest one managed alone.
  • On Terminal-Bench 4, built around punishing multi-step tasks, raw-trace fine-tuning reached 4.5% pass@3 from a 1.5% base, while the rewritten data reached 9.1%.
  • Rejection sampling in the rewrite loop expanded 2,001 source successes into 11,094 verified trajectories, all executed in standard bash.
  • The fine-tuned weights are public on Hugging Face as IntelligenceLab/RSR-27B.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Teams that bank successful runs from scaffolded harnesses for fine-tuning need a rewrite pass in the deployment harness first, since in this paper the raw traces left the model below its untuned score.
  • capability Heavy harnesses can stay in offline discovery while the shipped agent runs in a bare shell, avoiding the latency, state complexity and brittle assumptions the post attributes to deployed scaffolding.
  • cost Each kept trace costs a planner pass, at least one critic audit and a full sandboxed rerun with verification, on top of the discovery runs that found the original solution.
  • constraint Leak filtering is a model judgment: the critic is the same 27B base model checking runbooks against the public task description, so leak types it cannot recognise reach the training set.

Raw harness traces made Qwen-3.8-27B worse on Terminal-Bench 2. Trained directly on its own successful runs, the model finished 3.6 points below the untuned version's 57.0% pass@3 [9][20]. The paper is arXiv:2610.02826, from IntelligenceLab and the University of Maryland [2]. A dev.to post summarising it says the model learned the harnesses' surface artifacts instead of the terminal mechanics underneath [8]. Some traces leaned on StateM's explicit transition assertions. Others leaned on self-reflection prompts or leaked evaluation scripts into the bash history [6]. A run that solves a Linux task by reading the internal verifier script is a fine record of how to read verifier scripts. Train on it, the post argues, and the model expects that leak in production, then falls apart in a bare bash shell [24].

The heavy harnesses still found the solutions [4]. StateM handled structured multi-phase setup tasks that needed state kept across interruptions. The reflection harness handled trial-and-error debugging [5]. Recursive Self-Rewrite splits the job in two: discovery under whatever harness works, then execution under a vanilla one [19]. The same base model plays all three roles [13]. Each successful trace goes through the same loop:

1. The planner reads the trace and writes a runbook of milestones, state checks and recovery paths. Its instructions say to describe validation procedures without copying output deliverables or hardcoded solutions [10]. 2. The critic audits the runbook against the public task description. It rejects verifier leakage, direct answers and harness-specific syntax, and a rejected runbook goes back for another rewrite [11]. 3. The executor runs the task from scratch in a fresh, isolated sandbox under vanilla Terminus 2, with the runbook as private guidance [12]. 4. The runbook is stripped from the saved trajectory [12].

Step 4 is the part I would copy. The runbook shapes generation but never reaches the training record, so the student sees only bash actions that worked with no harness around them [12]. Terminus 2, the execution harness, was also one of the three discovery harnesses [3].

Against raw-trace training, the rewritten data gained 20.8 points on Terminal-Bench 2 [21]. The post says it beat both the base model and direct SFT on every benchmark tested [1]. Software Terminal-Bench 100 doubled from 3.0% to 6.0% pass@3 [16]. Process reward on Long-Horizon Terminal-Bench rose from 0.21 to 0.29 [17]. At 6.0%, the model still misses 94% of that suite's tasks across three attempts [22].

These are the authors' numbers on the authors' tasks. They describe your agent only if production looks like a general terminal loop and the work looks like multi-step engineering tasks, because every reported score comes from a Terminal-Bench variant [9][15][16][17]. The post does not say whether the roughly 3,000 training tasks overlap the evaluation suites [3]. Both training sets came from that pool [3][14]. Any overlap would affect both, so I would trust the gap between the two methods more than either absolute score.

What to watch

  • Independent Terminal-Bench 2 runs of the public RSR-27B weights that reproduce, or fail to reproduce, the 74.2% pass@3.
  • A published compute or sampling budget for the planner, critic and executor loop, to price each kept trajectory.
  • A rerun with a critic model different from the student, to test how many leaks the same-model critic lets through.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories