SDSignal Desk

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Oct 7, 2026, 9:00 AM · Microsoft Research

Image: Microsoft Research

Microsoft Research wants agents trained in the same harness they ship in. That closes a real gap between lab and deployment, and it makes reward design and safeguards matter even more.

Why it matters

Microsoft Research Asia has open-sourced Agent Lightning v1.0, a rebuilt framework for training AI agents with reinforcement learning. The pitch is in one idea the team calls Harnessed Agentic RL: whatever harness an agent uses in deployment is the same harness that takes part in training. Developers point their agent's model endpoint at an Agent Lightning proxy instead of rewriting the agent inside a training framework.

The whole thing is about 3,500 lines of code. Agents run as standard Kubernetes jobs rather than on paid commercial sandbox services. In the team's coding-agent example, built on SWE-smith, mini-SWE-agent and Qwen3.5-9B, reinforcement learning raised Pass@1 on SWE-bench Verified from 41.8% to 56.4% using about 6,000 training samples.

From the desk

We think the core insight here is right, and it is one practitioners have felt for a while. Modern agents are not just a model; they are a model wrapped in context management, tool protocols and execution logic. Most reinforcement learning setups ask teams to rebuild that wrapper inside the trainer, which is costly and means the agent being improved is not quite the agent people use. Training through the real harness removes that mismatch. If it holds up, it lowers the bar for smaller teams to tune agents on their own workflows.

The small footprint is part of the argument. A 3,500-line control plane is something one engineer can read and reason about, and running rollouts on ordinary Kubernetes instead of metered sandboxes keeps costs down and the pipeline reproducible. The researchers also report that their Collocated Async RL approach, where rollouts and model updates share the same GPUs, ran about twice as fast end to end as synchronous training while using fewer GPUs than conventional asynchronous setups. That is the kind of practical engineering that turns a paper idea into something people can actually run.

The benchmark gain is meaningful but should be read for what it is: one model, one harness, one benchmark. A 14.6-point jump on SWE-bench Verified from a 9B model with a modest dataset is a strong result. It does not tell us how well the recipe generalizes to other harnesses or messier real-world tasks.

The downside we want named is the flip side of making agent training easy. Reinforcement learning optimizes for whatever the reward measures, and agents are good at finding shortcuts. The team says its pipeline includes reward-hacking safeguards, which is reassuring, but as this approach spreads to teams with fewer resources, those safeguards will be the first thing to get skipped. Cheaper training of agents that act in real environments means more agents tuned hard toward narrow goals. I'm watching whether the open-source release makes good reward design and evaluation as easy to copy as the training loop itself.

Context

Earlier agent RL systems such as verl, AReaL and slime assumed the training framework owned the agent's loop. Agent Lightning v1.0 instead sits between the agent and the model as an OpenAI-compatible proxy that records prompts, responses and log probabilities. The researchers describe four problems that come with this design: retokenization, advantage calculation, loss normalization and scheduling variable workloads onto fixed GPUs.

Who feels it

Developers
Teams with existing agents can try reinforcement learning without rewriting them, and run rollouts on infrastructure they already have.
Enterprises
Tuning an in-house agent on company workflows becomes more realistic, provided teams invest in sound reward design and evaluation.
Researchers
A small, readable framework makes it easier to reproduce and extend agent RL experiments beyond a single benchmark.
Sandbox providers
Native Kubernetes support is a direct alternative to paid sandbox services for large rollouts.

What to watch

  1. Results from Agent Lightning on harnesses other than mini-SWE-agent, such as OpenHands or OpenCode
  2. Independent reproductions of the SWE-bench Verified gain
  3. How the project documents and shares reward-hacking safeguards
  4. Adoption by teams training agents on non-coding tasks

Read the original

Continue at the source.

Microsoft Research

Companies: Microsoft