On-policy distillation methods for LLMs
TECH

On-policy distillation methods for LLMs

32+
Signals

Strategic Overview

  • 01.
    On-policy distillation (OPD) is a post-training technique where a student model generates its own rollouts and a stronger teacher model grades every token of those trajectories with dense, token-level supervision, typically via reverse KL divergence.
  • 02.
    Because the student is graded on states it actually visits, OPD sidesteps the exposure-bias problem that plagues static off-policy distillation, where a model trained only on teacher-written text tends to compound errors once it starts generating on its own.
  • 03.
    The approach was popularized by Thinking Machines Lab's October 2025 blog post, which reported that OPD reached 74.4% on AIME'24 using just 1,800 GPU-hours, versus 67.6% for pure reinforcement learning at 17,920 GPU-hours and 60% for off-policy distillation trained on 400,000 examples.
  • 04.
    A dense wave of academic papers posted to arXiv in late September 2026 - including IPD, Interactive-Policy Distillation, SIPO, SAKI, and B-OPSD - refine specific failure modes of the original recipe, from teacher misalignment early in training to token-level credit assignment and prompt efficiency.

Why On-Policy Beats Both RL and Off-Policy Distillation

On-policy distillation departs from classic knowledge distillation by grading the student on trajectories the student itself generates, rather than on fixed text written by the teacher in advance. A stronger teacher model scores every token of the student's own rollout, typically using reverse KL divergence, which gives a dense, per-token reward signal similar to reinforcement learning but far cheaper to compute [1]. Because supervision happens on states the student actually visits, on-policy distillation avoids the exposure-bias problem that hurts purely off-policy approaches, where a model trained only on teacher-written text tends to drift and compound errors once it starts generating on its own [2]. In Thinking Machines Lab's original benchmark, this combination let a student reach 74.4 percent on AIME'24 using only 1,800 GPU-hours, versus 67.6 percent for pure reinforcement learning at 17,920 GPU-hours and 60 percent for off-policy distillation trained on 400,000 examples - a claimed 7 to 10x speed advantage over RL and a 50 to 100x reduction in cumulative compute [1].

The September 2026 Wave: One Recipe, Five Different Fixes

A year after that original post, late September 2026 produced a dense cluster of academic papers that each target a specific crack in the base recipe rather than proposing a wholesale replacement. Interpolated Policy Distillation (IPD) turns the off-policy/on-policy split into a tunable continuum, blending student and teacher token distributions token-by-token with speculative decoding to keep the extra compute manageable [3]. Interactive-Policy Distillation addresses what researchers call teacher unanchoring - early in training a weak student wanders into prefixes far from anything the teacher has seen - by having student and teacher alternate propose-and-verify roles, reporting meaningful accuracy gains using roughly a quarter of the training examples of standard on-policy distillation [4]. SIPO folds reinforcement learning and self-distillation into one contrastive self-teacher, pairing correct reference answers with the student's own mistakes for dense token-level credit assignment without needing an external teacher at all [5]. SAKI, from a Meituan and KTH collaboration, routes supervision between reverse-KL and direct token correction depending on where the student's rollout diverges from the teacher, and reports a large rollout-throughput improvement via a speculative verifier [6]. B-OPSD tackles the anchoring problem differently, temporarily training the policy ahead of itself to create a 'future teacher' before resetting and supervising the original student against it [7]. Other papers in the same window target orthogonal weaknesses: one shows that smaller students distilled in thinking mode fail to inherit the teacher's stopping behavior and keep generating past a correct answer even though they still learn to solve the problem [8]; another finds that prompt usefulness for distillation is highly uneven and relational to the specific teacher rather than an intrinsic property of the prompt - a handful of prompts can match the performance of thousands, but deliberately curating which ones does not reliably beat picking them at random [9]; and a third introduces per-token weighting via bilevel optimization, letting a smaller student actually surpass its own larger teacher on math benchmarks [10].

The Compute, Memory, and Tokenizer Tax

None of this comes for free. On-policy distillation requires the teacher model to stay resident in memory and run inference in parallel with the student for the full duration of training, a structural cost that off-policy approaches simply do not carry. That is part of why one practitioner reflection frames on-policy distillation as a tool that only pays off when off-policy data genuinely fails to cover the states the student will actually visit, and works best when the student and teacher are reasonably well matched in capability and the task does not require very long generations [11]. Most published variants, including the original recipe, are also largely restricted to teacher-student pairs from the same model family so that token-level supervision lines up cleanly - a constraint that has pushed a separate strand of community work toward extending the technique across different tokenizers, since real deployments rarely get to pick a same-family teacher for free.

Knowledge Transfer or Just Token Suppression?

The efficiency numbers have not settled the question of what on-policy distillation is actually teaching the student. Some observers read the mechanism as straightforward knowledge transfer: the teacher's per-token judgments propagate genuine reasoning ability into the student's weights. Others, pointing to how the reverse-KL-style reward term behaves in practice, argue the visible gains lean more toward suppressing the student's bad tokens - a bounded, mostly negative correction - than toward injecting new positive knowledge from the teacher. A separate line of research lends some support to the suppression reading: rather than grading every token, one recent paper found that concentrating the teacher's supervision on as little as one percent of a trajectory's tokens, and sometimes as little as a tenth of that, can match the results of scoring every token, which suggests only a small fraction of tokens are actually carrying the corrective signal. That distinction matters for how far the technique can scale: if the real mechanism is closer to pruning bad behavior than transferring capability, the eye-catching efficiency numbers may not extrapolate cleanly to harder tasks or bigger capability gaps between teacher and student. The debate has not stayed confined to papers and forum threads, either - within a year of the original announcement, the topic had spread to standalone explainer videos and a podcast segment walking through the mechanism for a broader technical audience, a sign the idea has moved past a narrow research niche.

Historical Context

2025-10-27
Published the 'On-Policy Distillation' blog post, popularizing the technique for LLM post-training and reporting major compute-efficiency gains over reinforcement learning.
2026-04-01
Published 'A Survey of On-Policy Distillation for Large Language Models,' formalizing a taxonomy across three design axes: what to optimize, where the signal comes from, and how to stabilize training.
2026-09-29
A cluster of five or more independent papers - IPD, Interactive-Policy Distillation, SIPO, SAKI, and B-OPSD - were posted to arXiv within the same short window, each targeting a different weakness in the original recipe.

Power Map

Key Players
Subject

On-policy distillation methods for LLMs

TH

Thinking Machines Lab

Published the widely-cited October 2025 blog post 'On-Policy Distillation' that popularized the technique for LLM post-training, applying it to math reasoning and an internal chat assistant and reporting the headline compute-efficiency numbers that still anchor the field's framing a year later.

QW

Qwen team (Alibaba)

Cited as an inspiration for Thinking Machines' work; Qwen3 reportedly used on-policy or strong-to-weak distillation in production post-training recipes.

AC

Academic authors of the September 2026 arXiv wave (including Meituan and KTH Royal Institute of Technology on SAKI)

Multiple independent research groups published refinements to on-policy distillation, including IPD, Interactive-Policy Distillation, SIPO, and SAKI, within days of each other.

MI

Mingyang Song and Mao Zheng

Authors of 'A Survey of On-Policy Distillation for Large Language Models,' formalizing a taxonomy across three design axes - what to optimize, where the signal comes from, and how to stabilize training - several months before the September 2026 paper cluster.

Fact Check

11 cited
  1. [1] On-Policy Distillation (Thinking Machines Lab blog)
  2. [2] On-Policy Distillation (EmergentMind)
  3. [3] Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
  4. [4] Interactive-Policy Distillation with Bidirectional Propose-and-Verify
  5. [5] SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
  6. [6] SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
  7. [7] Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
  8. [8] Solving Without Stopping: On-Policy Distillation at Small Scale
  9. [9] Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
  10. [10] Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
  11. [11] On-Policy Distillation: A Reflection

Source Articles

Top 5

THE SIGNAL.

Analysts

“Describes on-policy distillation as mechanically similar to DAGGER from robotics and imitation learning, with the key difference that on-policy distillation for LLMs typically works best with reverse KL (or Jeffreys divergence, half forward and half reverse) rather than the forward-KL or behavior-cloning losses typical in robotics and RL.”

Rishabh Agarwal
ML researcher

“Has publicly criticized on-policy distillation, arguing the teacher model is placed in what he calls a 'structurally terrible position' during the process.”

Omar Khattab
Creator of DSPy, MIT
The Crowd

“Our latest post explores on-policy distillation, a training approach that unites the error-correcting relevance of RL with the reward density of SFT. When training it for math reasoning and as an internal chat assistant, we find that on-policy distillation can outperform other https://t.co/ltPsdNjajD”

@@thinkymachines2775

“colab notebook for on-policy distillation 👇🔗 (for those without @thinkymachines tinker access) train qwen-0.6b with OPD to get from 38% -> 60% on GSM8K works for models without the same tokenizer!”

@@itsandrewgao292

“Can supervising just 1% of tokens match full OPD? Our new paper shows it can, and sometimes, 0.1% is enough! 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation”

@@HuanxinShe525483

“On-policy distillation: one of the hottest terms on PapersWithCode [R]”

@u/NielsRogge105
Broadcast
How Small Models Learn to Think Like Giants | On-Policy Distillation

How Small Models Learn to Think Like Giants | On-Policy Distillation

How On Policy Self Distillation Works - Sasha Rush

How On Policy Self Distillation Works - Sasha Rush

On Policy Distillation - How the big AI labs actually train their LLMs

On Policy Distillation - How the big AI labs actually train their LLMs