Microsoft Research has released Agent Lightning v1.0, an open-source framework that trains AI agents with reinforcement learning (RL) while leaving the agent's real harness untouched. The pitch is simple: stop reimplementing your agent inside a training framework, and let the same tools, context management and control flow you ship in production generate the training data. In the authors' coding-agent run, that took Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified, a 14.6 point gain from roughly 6,000 training examples. The framework is about 3,500 lines of code, MIT licensed, and available on GitHub.
If you build agents, the interesting part is not the leaderboard number. It is the claim that the harness, the layer that decides which tools an agent sees and how its context is packed, should be part of the training loop. Our guide to agent harnesses explains why that layer now matters as much as the model, and Agent Lightning is a bet that you should train through it, not around it.
TL;DR: Agent Lightning v1.0 at a glance
| Question | Answer |
|---|---|
| Who built it? | Microsoft Research Asia, described in the Microsoft Research post |
| Is there a paper? | Yes, arXiv 2608.17528, submitted August 18, 2026 |
| How big is it? | About 3,500 lines of code |
| License? | MIT |
| Headline result? | Qwen3.5-9B: 41.8% to 56.4% on SWE-bench Verified with about 6K examples |
| Newer example? | Qwen3.5-35B-A3B from 47.8% to 61.6% on SWE-bench Verified (September 2026 repo note) |
| Where do agents run? | As Kubernetes Jobs, no paid sandbox service required |
| Training backend? | Built on verl |
What problem does harnessed agentic RL solve?

Classic agentic RL frameworks run the environment loop themselves. You write the agent a second time inside the trainer, with its own prompt format, its own tool wrappers and its own idea of when to stop. The result is a gap: the policy you trained was never exposed to the exact harness that serves users.
The paper's abstract frames the alternative. "Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system." In harnessed agentic RL, the harness owns the interaction loop and the trainer "observes only sequences of LLM request-response pairs." The original Agent Lightning introduced that disaggregated design through an LLM endpoint proxy, and the authors say later frameworks such as verl Uni-Agent, AReaL 2.0, slime and Polar adopted the same paradigm.
For practitioners this means the agent code changes very little. You point the harness at the Agent Lightning gateway instead of a model provider, run tasks, and attach a reward. Existing harnesses such as mini-SWE-agent and OpenHands are the examples Microsoft names. If you have been following the shift toward harness engineering, see how LangChain tuned a harness against Terminal Bench for the evaluation-side version of the same idea.
How does Agent Lightning work?
The framework has three pieces, per the technical report.
- API Gateway. A stateless service that stores rollouts, models and events, and exposes an OpenAI-compatible proxy endpoint that the harness calls.
- Rollout Controller. Launches agent runs as Kubernetes Jobs, or as a local process pool, using a standard reconciliation pattern.
- Customized Trainer. Built on verl, it turns recorded calls into training samples and handles the algorithmic details described below.
Because agents run as ordinary containers, you can reuse the images and task environments you already use for evaluation. Microsoft highlights that this avoids depending on commercial sandbox services. If you are weighing how to isolate agent runs generally, our notes on Cloudflare's Sandbox SDK 1.0 cover the hosted alternative.
The hard parts the paper names
This is where the post earns its keep, because the paper is candid that harnessed RL is harder than it looks. The harness talks in text, but RL updates tokens. That mismatch causes several problems.
Retokenization breaks prefixes
When a harness sends the conversation back to the model each turn, the text is tokenized again. The authors identify three causes of drift: chat templates that are not compositional, decode-then-retokenize differences, and harness-side output transformations. If the new prompt tokens do not exactly extend the previous ones, the calls cannot be treated as one continuous sequence.
Agent Lightning uses best-effort sequence merging. It merges consecutive calls only when the token IDs match exactly, which preserves the prompts the model really saw at deployment. The price is fewer merges. In coding tasks the paper reports an average of 36% single-sample rollouts, meaning many rollouts end up as several training samples.
Variable sample counts distort the gradient
A single rollout can produce a different number of samples depending on retokenization, subagent spawning or context summarization. If you normalize at the sample level, rollouts that happen to split more get more weight for incidental reasons. The paper argues for rollout-level advantages and a rollout-level token-mean loss so that each rollout counts equally. In the authors' words, sample count "should not be allowed to affect gradient normalization."
Scheduling gets messy
Dynamic sample counts make fixed GPU allocation awkward. The system keeps rollout membership intact and avoids splitting one rollout across optimizer updates, which would mix policy versions inside a single trajectory.
Collocated asynchronous RL
To use hardware efficiently, Agent Lightning shares the same GPUs between generating rollouts and updating the model. The report claims roughly a 2x end-to-end speedup over synchronous RL while using fewer GPUs than a dedicated asynchronous setup. This is the authors' measurement on their stack, so expect your numbers to depend on your cluster and workloads.
What do the results actually show?
| Agent type | Model | Before | After |
|---|---|---|---|
| Coding (SWE-bench Verified) | Qwen3.5-9B | 41.8% | 56.4% |
| Search (validation reward) | Llama-3.2-3B | 25.1% | 41.7% |
| Instruction following | Qwen3-4B | 51.9% | 70.2% |
The coding run used about 6,000 training examples, and the paper describes the compute as "modest." The authors also release the complete workflow and training scripts, which is the part most worth trying: a reproducible coding-agent RL pipeline is rare.
A few cautions keep this honest. These numbers come from the authors, not an independent lab. SWE-bench Verified has known quirks, and our coverage of OpenAI's audit of broken SWE-bench Pro tasks and Cursor's reward-hacking findings is a reminder to check what a benchmark gain really means. When you reward an agent through RL, reward hacking is a live risk, so inspect trajectories, not only scores.
Why does this matter for builders?
Small models can be tuned for your harness. A 9B model moving 14.6 points suggests that harness-specific post-training can close a meaningful part of the gap to larger models on narrow work. That is attractive for teams that want cheaper, local or private agents.
It treats the harness as part of the product. Teams already iterate on prompts, tools and context packing. Training through the same harness means improvements in one carry over to the other instead of being lost in translation. This connects to the thread in Self-Harness: agents that improve themselves, where the harness itself is the thing being optimized.
Reliability still needs measurement. Training lifts averages, but production needs consistency. Microsoft's separate ThinkingBox benchmark found that even top models pass all 20 runs on fewer than half of tasks, which is the kind of check you should pair with any RL-tuned agent.
It lowers the engineering bar. At about 3,500 lines the codebase is small enough to read in an afternoon. That matters for research teams who want to modify advantage calculation or merging rules instead of fighting a large framework. For a broader view of the harness landscape, see our top 10 open and closed source agent harnesses.
How do I try it?
The repository's setup uses uv. Per the README, you clone the repo, run uv sync, then run bash scripts/setup_verl.sh 0.8.0 cu130 to set up verl. The installation guide and documentation cover the rest. A realistic path:
- Pick a harness you already run, such as mini-SWE-agent.
- Point its model endpoint at the Agent Lightning gateway.
- Start with the released coding-agent workflow on a small model to confirm the pipeline reproduces.
- Swap in your own tasks and a reward you trust, and watch trajectories for shortcuts.
You will need GPUs and a Kubernetes cluster for the full design. As of the repository snapshot we checked, the project had roughly 18,600 stars, and its README reports the MoE example improving Qwen3.5-35B-A3B from 47.8% to 61.6% on SWE-bench Verified after training on only 1.8K examples.
Questions people are asking
Is this the same as fine-tuning on transcripts? No. Supervised fine-tuning copies recorded behavior. RL rewards outcomes, so the model can discover better strategies inside the harness.
Can I train a closed API model with it? The method needs access to model weights for updates, so it targets open-weight models you can host.
Does it replace evaluation? No. It produces a trained policy; you still need held-out tasks and repeated runs to know whether the gain holds.
Bottom line
Agent Lightning v1.0 is a compact, reproducible answer to a real gap in agent training: the model you train should see the harness you ship. The reported gains are large for small models, the design avoids paid sandboxes, and the paper is unusually open about retokenization and normalization pitfalls. Treat the numbers as the authors' results, read the trajectories, and use the released pipeline as a starting point.
Version numbers, benchmark figures and repository details are accurate as of October 7, 2026 and may change as the project evolves.
Related reading
- What is an agent harness? The complete guide
- Agent harness engineering with Terminal Bench and LangChain
- Self-Harness: agents that improve themselves
- Top 10 open and closed source agent harnesses
- Microsoft ThinkingBox: grading agents on database state
- Cursor reward hacking and SWE-bench contamination
- Official docs: Agent Lightning
