Signum
Feed
Useful signal7 Oct 2026high confidence

Microsoft Research Asia open-sources Agent Lightning v1.0, a ~3,500-line framework for reinforcement learning on agents using their real deployment harnesses

Microsoft Research Asia released a fully rebuilt Agent Lightning v1.0 as open source. It places an OpenAI-compatible LLM proxy between an existing agent harness and the model, so the harness can be trained without being reimplemented in the training framework. The framework has three components (API Gateway, Rollout Controller, Customized Trainer built on verl). It runs agents as local processes or Kubernetes jobs, and adds Collocated Async RL, where rollout and training share the same GPUs. The authors report an example coding-agent pipeline that raised Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified (+14.6 points) with about 6,000 training samples. The results are self-reported.

CapabilityInfrastructureAdoption

Entities: Microsoft Research Asia, Microsoft Research, Agent Lightning, Qwen3.5-9B, SWE-bench Verified, verl

66Useful signal
1 source
1 primary
Was this useful?
01

What happened

Microsoft Research Asia released Agent Lightning v1.0 as open source, a rebuilt version of an existing framework of about 3,500 lines. It puts an OpenAI-compatible proxy between an agent's existing harness and the model, so the harness can be trained as it is deployed rather than reimplemented inside a training framework. It has three parts (API Gateway, Rollout Controller, and a trainer built on verl), runs agents as local processes or Kubernetes jobs, and lets rollout and training share the same GPUs. The authors report that an example coding-agent pipeline lifted Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified (+14.6 points) using about 6,000 training samples. These figures are self-reported.

02

Why it matters

The main audience is teams that already run an agent harness and want to improve an open-weight model on it with reinforcement learning. For them, the proxy approach could cut the engineering work of rebuilding the harness for training and avoid the mismatch between how an agent is trained and how it is deployed. Running on Kubernetes may also remove the need for paid commercial sandbox services. The practical impact is limited to groups with GPU capacity and RL expertise, and the benefit beyond this one example is unproven.

03

What is noise

The headline gain comes from one example pipeline on a small 9B model, reported by the authors, so it says little about how well the method generalises to other harnesses or larger models. The small line count is a neat selling point but says nothing about quality or ease of use. This is also a rebuilt v1.0 of an existing project, so the "new paradigm" framing overstates the novelty, and the extraction found no primary evidence links to check.

04

Watch next

  1. 01Independent reproduction of the 41.8% to 56.4% SWE-bench Verified result, ideally with other harnesses such as OpenHands or mini-SWE-agent and other base models, within the next 1 to 3 months.
  2. 02Adoption signals on the GitHub repo: external contributors, issues about real deployments, and any third-party projects or companies reporting they trained a production harness with it.
  3. 03Whether the gains hold up against contamination and overfitting checks, and whether the same method works at larger model sizes or with closed-model harnesses like Claude Code or Codex, including reported compute costs for the Collocated Async RL setup.

Coverage

1 story

More capability signals

Full feed →