Science1 publisherNot yet confirmed elsewhere2 min readPublished
Microsoft's Agent Lightning trains AI agents inside the harness they ship with
Microsoft Research Asia open-sourced Agent Lightning v1.0, a 3,500-line framework whose recipe lifted a 9B model's SWE-bench Verified score by 14.6 points. It runs reinforcement learning on the agent a team already deploys, sitting as a proxy between that agent and its model.
The Scientist · Science desk
What happened
- The coding recipe used only about 6,000 training samples, built from a dataset that is already open source.
- Rollouts run as standard Kubernetes jobs on self-managed, cloud or local clusters, with no paid commercial sandbox service required.
- Because the trainer sees only request and response pairs, one rollout can arrive in a variable number of parts, one of four challenges the post lists.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- capability A team can put its production agent, unchanged, into an RL loop by redirecting one model endpoint, so whatever training teaches is learned by the code that ships.
- cost Running rollouts on a team's own Kubernetes cluster swaps a paid sandbox service for cluster capacity the team already operates and pays for.
- decision With the harness and sandbox out of the way, a task-specific fine-tune now turns on whether a team can score its agent's outcomes, since RL learns only from the rewards it is given.
The coding result is a before-and-after on one model and one benchmark, and the figures are Microsoft's own. The effect is sizeable. At 41.8% Pass@1, Qwen3.5-9B left 58.2% of SWE-bench Verified tasks unsolved, and after training it left 43.6% [8][12]. Training removed about a quarter of the model's failures [13].
The design is what makes the result interesting. RL systems such as verl, AReaL and slime were built on the assumption that the trainer owns the agent's loop: the model acts, the environment returns an observation, the observation joins the context, and the rollout becomes one continuous token trajectory [1]. Production agents carry their own machinery. OpenHands, mini-SWE-agent and OpenCode each bring their own context management, tool protocols and execution logic [2]. Rebuilding one inside a trainer is expensive, Microsoft argues, and the copy may not behave like the original [3].
Agent Lightning leaves the loop where it is. It sits between agent and model as a proxy. The team points the agent's model endpoint at it, and it records each call while the harness code stays unchanged [10].
Not owning the loop moves the difficulty to the trainer. It sees only a series of request and response pairs, so a single rollout may arrive split into a variable number of parts, one of four challenges the post sets out for training on real harnesses [11]. I think that is the right place for the difficulty. It gets solved once, in a codebase of about 3,500 lines that a team can read and change [6], instead of in every team's reimplementation of its own agent.
For an in-house fine-tune, the post takes two items off the bill: the rebuild and the commercial sandbox [7][3]. The data requirement is small too, about 6,000 samples from an open-sourced dataset [9]. The post does not report the GPU hours or hardware behind the run, or how the same model would score if trained on a rebuilt loop with the same data. In my view, compute is the line that decides whether a small task-specific model pays for itself, and it is the one still unpriced. The missing comparison matters for a different reason. The experiment shows this pipeline improves the model; the claim that training on the real harness is the cause rests on the design argument.
What to watch
- An independent reproduction of the 41.8% to 56.4% result, or the same recipe run on a second model or a different harness.
- Published GPU hours and hardware for the Qwen3.5-9B coding run, which would let a team price an in-house fine-tune.
- An ablation training the same model on the same data through a rebuilt loop, to test whether the real harness is the cause of the gain.