Skip to content

Science1 publisherNot yet confirmed elsewhere2 min readPublished

Microsoft's Agent Lightning trains AI agents inside the harness they ship with

Microsoft Research Asia open-sourced Agent Lightning v1.0, a 3,500-line framework whose recipe lifted a 9B model's SWE-bench Verified score by 14.6 points. It runs reinforcement learning on the agent a team already deploys, sitting as a proxy between that agent and its model.

The Scientist · Science desk

How we use AISend a correction

What happened

  • The coding recipe used only about 6,000 training samples, built from a dataset that is already open source.
  • Rollouts run as standard Kubernetes jobs on self-managed, cloud or local clusters, with no paid commercial sandbox service required.
  • Because the trainer sees only request and response pairs, one rollout can arrive in a variable number of parts, one of four challenges the post lists.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • capability A team can put its production agent, unchanged, into an RL loop by redirecting one model endpoint, so whatever training teaches is learned by the code that ships.
  • cost Running rollouts on a team's own Kubernetes cluster swaps a paid sandbox service for cluster capacity the team already operates and pays for.
  • decision With the harness and sandbox out of the way, a task-specific fine-tune now turns on whether a team can score its agent's outcomes, since RL learns only from the rewards it is given.

The coding result is a before-and-after on one model and one benchmark, and the figures are Microsoft's own. The effect is sizeable. At 41.8% Pass@1, Qwen3.5-9B left 58.2% of SWE-bench Verified tasks unsolved, and after training it left 43.6% [8][12]. Training removed about a quarter of the model's failures [13].

The design is what makes the result interesting. RL systems such as verl, AReaL and slime were built on the assumption that the trainer owns the agent's loop: the model acts, the environment returns an observation, the observation joins the context, and the rollout becomes one continuous token trajectory [1]. Production agents carry their own machinery. OpenHands, mini-SWE-agent and OpenCode each bring their own context management, tool protocols and execution logic [2]. Rebuilding one inside a trainer is expensive, Microsoft argues, and the copy may not behave like the original [3].

Agent Lightning leaves the loop where it is. It sits between agent and model as a proxy. The team points the agent's model endpoint at it, and it records each call while the harness code stays unchanged [10].

Not owning the loop moves the difficulty to the trainer. It sees only a series of request and response pairs, so a single rollout may arrive split into a variable number of parts, one of four challenges the post sets out for training on real harnesses [11]. I think that is the right place for the difficulty. It gets solved once, in a codebase of about 3,500 lines that a team can read and change [6], instead of in every team's reimplementation of its own agent.

For an in-house fine-tune, the post takes two items off the bill: the rebuild and the commercial sandbox [7][3]. The data requirement is small too, about 6,000 samples from an open-sourced dataset [9]. The post does not report the GPU hours or hardware behind the run, or how the same model would score if trained on a rebuilt loop with the same data. In my view, compute is the line that decides whether a small task-specific model pays for itself, and it is the one still unpriced. The missing comparison matters for a different reason. The experiment shows this pipeline improves the model; the claim that training on the real harness is the cause rests on the design argument.

What to watch

  • An independent reproduction of the 41.8% to 56.4% result, or the same recipe run on a second model or a different harness.
  • Published GPU hours and hardware for the Qwen3.5-9B coding run, which would let a team price an in-house fine-tune.
  • An ablation training the same model on the same data through a rebuilt loop, to test whether the real harness is the cause of the gain.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories