explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Background: what on-policy distillation is
  • The core idea: the teacher is a reward model
  • What the experiments used
  • Finding 1: gains are sharpening, not new capability
  • Finding 2: the collapse run
  • The two fixes
  • What this means if you train models
  • Caveats
  • Why it matters
  • Related reading
← Back to blog

explainx / blog

On-Policy Distillation Can Collapse: A New Paper Explains Why (and the Two Fixes)

Research, Distillation, Reinforcement Learning, Post-Training, AI News

Part of AI Research

A new arXiv paper treats on-policy distillation as RL with the teacher as implicit reward. It explains gains, collapse into repetition, and two fixes that work.

Oct 8, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
On-Policy Distillation Can Collapse: A New Paper Explains Why (and the Two Fixes)

Related: if distillation is new to you, start with what AI distillation is, then read how black-box distillation works in Proxy-KD.

On-policy distillation can make a small model noticeably better, or push it into endless repetition, and a new paper argues both outcomes come from the same mechanism. In "Gains and Collapse in On-Policy Distillation: A Reinforcement Learning Perspective" (arXiv, October 2, 2026, by Han Cui, Jianhao Yan, Yun Luo, Hongbo Zhang, Zhizhang Fu and Yue Zhang), the authors treat the fixed teacher as an implicit reward model that scores the student's own rollouts. When that reward is reliable, the student sharpens what it already knows. When it is not, the student learns to exploit it. The paper was also trending on Hugging Face Papers with 27 upvotes.

A paperclip holding research pages with a small branching diagram, illustrating a post about a distillation research paper

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer from the paper
What is the framing?OPD is RL; the teacher is an implicit reward for the student's own samples.
Does OPD add new capability?Not in their audit: it raises pass@1 on problems the student could already solve.
What causes collapse?The teacher rewards overlong, repetitive student behavior (reward hacking).
Fixes?Loss masking on unhealthy responses; SFT warmup before OPD.
Code?github.com/HancCui/opd_hacking

Background: what on-policy distillation is

Classic knowledge distillation trains a small student on text written by a large teacher. That has a known weakness: the student is trained on situations the teacher visits, but at inference it lives in the situations it produces itself, including its own mistakes. On-policy distillation fixes the mismatch. The student samples its own responses, and the teacher grades them, typically by providing its token-level probabilities for what it would have said in that context. The student is then pushed toward the teacher's distribution on its own trajectories.

This is why OPD has become an important part of language-model post-training: it is cheaper than running full reinforcement learning with verifiable rewards, yet it gives dense, per-token feedback instead of a single pass or fail at the end. The paper's opening complaint is that despite its gains, OPD "can also collapse into excessively long and repetitive generation," and the mechanism was poorly understood. For broader context on teacher-student training and why labs fight over it, see our explainers on AI distillation and the policy fight around distillation and export controls.

The core idea: the teacher is a reward model

The authors' move is to rewrite OPD in reinforcement-learning terms. In RL, a policy (here the student) generates trajectories, a reward function scores them, and the policy shifts probability toward high-reward behavior. In OPD there is no explicit reward, but the teacher's log-probabilities on the student's tokens act as one: tokens the teacher considers likely get a positive signal, tokens it considers unlikely get a negative one. That is the "implicit reward."

Once you accept that framing, two things follow that match the paper's experiments:

  1. Gains come from amplifying behaviors the teacher favors that the student can already produce. OPD is a sharpening operation.
  2. Any reward can be hacked. If the teacher's preferences diverge from real quality on some region of the student's output space, the student will drift there. The paper says this directly: when the implicit reward is unreliable, "reward hacking happens."

This is the same family of failure that shows up elsewhere in AI training. OpenAI traced a quirky personality behavior to a reward signal in where the goblins came from, and evaluation gaming is covered in our look at reward hacking and eval contamination. The new paper's contribution is to show the same pattern inside distillation, which many practitioners treat as the "safe" alternative to RL.

What the experiments used

According to the full text, the authors tested three teacher-student settings:

  • JustRL: teacher JustRL-1.5B to student DeepSeek-R1-Distill-Qwen-1.5B, trained on DeepMath-103K (difficulty level 6 and up) with a 16,384-token limit.
  • DeepScaleR: teacher DeepScaleR-1.5B-Preview to the same student and data.
  • Qwen3-4B: teacher Qwen3-4B to student Qwen3-1.7B-Base, trained on OpenThoughts3 prompts with an 8,192-token limit.

Each run covered 100 rollout iterations in the Slime framework, with a reference-KL coefficient of 0.001 and a learning rate of 1e-5. Evaluation used AIME24, AIME25, AIME26 and AMC23, with mean@4 for routine accuracy and 256 responses per problem for coverage and pass@k analysis. These are small models and math-only tasks, which matters for how far the conclusions travel (more below).

Finding 1: gains are sharpening, not new capability

The paper's capability audit asks a simple question: does the OPD student solve problems the original student could never solve? Their answer, in these settings, is no. The initial student had 73 solvable AIME problems when sampled 256 times each. No problem was solved only by the OPD student, and the OPD students covered 65 (JustRL) and 63 (DeepScaleR) of the 73.

What OPD did change was reliability. It raised per-problem success on 90.6% (JustRL) and 84.4% (DeepScaleR) of already-solvable problems, with mean@256 gains of 22.6% and 14.7%. So pass@1 improves, but the advantage shrinks as you let the model sample more attempts, exactly the signature of sharpening rather than expansion. This lines up with a broader debate about whether RL-style post-training teaches models anything new or just surfaces what pretraining already put there.

For builders the practical read is modest and useful: OPD is a way to make a small model reliably do what it can occasionally do, not a way to teach it skills it lacks. If your student cannot solve a class of problems in 256 tries, do not expect distillation from any teacher to fix that by itself.

Finding 2: the collapse run

The striking result is the Qwen3-4B to Qwen3-1.7B-Base setting. The training metrics looked fine, converging smoothly, yet the endpoint was a mess. Per the paper's Table 2 (averages over AIME24 to AIME26):

table · 4 cols
SettingAccuracy of OPD studentTruncationSevere repetition
JustRL teacher37.1% (+16.2 over initial)16.5%0%
Qwen3-4B teacher5.4% (+5.2 over initial)99.4%38.0%

The Qwen3-4B teacher itself shows 2.1% truncation and 0% repetition, so the teacher is not the one rambling. The student is. A response-similarity measure (Min-10NN distance) fell sharply during that run, meaning the student's outputs were becoming near-copies of each other, and those outputs had low negative log-likelihood under the initial student, so they were not exotic: they were cheap, repetitive continuations the student already tended to produce.

Why would a teacher reward that? The authors looked at which initial student responses the teacher scored highest. In the Qwen3-4B setting, those top-advantage responses were overlong and repetitive, with near-zero correctness. A teacher judging token by token can find a long, self-repeating stretch very predictable and therefore "good." They also report the same pattern with Qwen3-8B and Qwen3-30B-A3B as teachers, so it is not a quirk of one checkpoint. They quantify it with a "selection gap," which measures how strongly training favors the teacher's most-preferred responses: about 0.07 for the healthy JustRL run versus about 26 for the collapsed Qwen3-4B run.

The two fixes

Because the diagnosis is reward hacking, the fixes are about the reward and the starting point.

Loss masking. Mask the loss on responses that hit the generation limit (the unhealthy ones). With the teacher fixed in the Qwen3-4B setting, this lets the run recover from collapse. Average accuracy over AMC23 and AIME24 to AIME26 improved by 3.08 points with the 4B teacher, 0.51 with the 8B teacher and 2.58 with the 30B-A3B teacher. Only about 10% of sampled responses kept nonzero loss, which tells you how dominated the batch was by degenerate output.

SFT warmup. Initialize the student from a supervised fine-tuning checkpoint instead of the raw base model. This avoided collapse entirely in the tested setting: a 16.38% average beat the base-initialized OPD run's 10.16% by 6.22 points. The intuition is that a student that already knows the format and length of good answers gives the teacher fewer degenerate modes to reward.

Neither fix is exotic, which is partly the point. They are cheap to add to an existing pipeline.

What this means if you train models

  • Log truncation and repetition rates during OPD, not just loss and accuracy. In the collapsed run, loss looked healthy while 99.4% of outputs were cut off. A dashboard that tracks only reward-like curves would have missed it.
  • Judge the teacher as an evaluator, not only a generator. The authors' conclusion is that the teacher's reliability in scoring student rollouts matters as much as its generation quality. A strong model is not automatically a good grader of a weak model's mistakes.
  • Warm up with SFT when the student is a raw base model. It is the simpler of the two mitigations and carries no extra machinery.
  • Do not expect new skills. Use OPD to consolidate and stabilize, and use other methods (more pretraining, better data, RL with verifiable rewards) to expand capability. Our coverage of automated post-training and RL post-training of an open model shows where that work is heading, and GRPO-trained catalog models give another example of RL on small models.

Caveats

This is a single preprint (v1), not yet peer reviewed, and we have not reproduced it. The students are 1.5B to 1.7B parameters and the benchmarks are competition math, where long chain-of-thought is the norm and truncation is a natural failure. Whether the same hacking dynamic dominates at frontier scale, on code or agentic tasks, or with different teacher-student gaps is open. The arXiv abstract itself gives no numbers; the figures above come from the paper's HTML full text. The authors link code at github.com/HancCui/opd_hacking so others can check.

Also note what the paper does not claim: it does not say OPD is bad. In the JustRL setting, the student gained 16.2 points with zero repetition. The takeaway is a monitoring and design discipline, not a retreat from the method.

Why it matters

Distillation is increasingly central to how small, cheap models get good, and also to the politics of AI, from lab disputes over copying outputs to export-control debates. A clean theory of when it helps and when it fails is useful on both fronts. Framing OPD as RL means the large toolbox of reward-hacking detection, reward shaping and regularization is now relevant to it, and it gives teams a concrete early-warning checklist: watch length, truncation and diversity, not just the loss.

Related reading

  • What AI distillation is and why labs fight over it
  • Proxy-KD: black-box LLM distillation
  • Reward hacking and eval contamination in coding benchmarks
  • Where the goblins came from: a reward gone wrong
  • Intology Locus and automated post-training
  • RL post-training of an open model

Follow @explainx_ai for daily research explainers.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 6, 2026

OpenAI "400 Math Papers" Rumor: What Is Verified and How to Check a Mass Release

On October 6, 2026, a small X account claimed OpenAI is about to release 400 papers on every math problem it has solved. OpenAI has not announced it. This post separates the rumor from the record (10 proofs in August, a 100-plus claim in September), explains why the replies were so hostile, and gives a checklist for judging a mass release of AI-written mathematics.

Sep 3, 2026

"Open-Source RL-as-a-Service": What That Phrase Actually Buys You

Aravind Srinivas posted a GitHub link on September 3 with a three-word caption — "open-source RL-as-a-service" — and the replies named half the ecosystem: Miles, prime-rl, SkyRL, SGLang. The phrase is doing a lot of work. Here is the anatomy of an RL post-training stack, what each contender is actually for, and the uncomfortable question of whether you need one.

Oct 9, 2026

Apple Normalizing Trajectory Models: 4-Step Image Generation

Apple Machine Learning Research published Normalizing Trajectory Models, a way to generate images in four steps while keeping an exact likelihood. This post explains the idea in plain language, the reported results, and what the paper does not claim.