Why On-Policy Beats Both RL and Off-Policy Distillation
On-policy distillation departs from classic knowledge distillation by grading the student on trajectories the student itself generates, rather than on fixed text written by the teacher in advance. A stronger teacher model scores every token of the student's own rollout, typically using reverse KL divergence, which gives a dense, per-token reward signal similar to reinforcement learning but far cheaper to compute [1]. Because supervision happens on states the student actually visits, on-policy distillation avoids the exposure-bias problem that hurts purely off-policy approaches, where a model trained only on teacher-written text tends to drift and compound errors once it starts generating on its own [2]. In Thinking Machines Lab's original benchmark, this combination let a student reach 74.4 percent on AIME'24 using only 1,800 GPU-hours, versus 67.6 percent for pure reinforcement learning at 17,920 GPU-hours and 60 percent for off-policy distillation trained on 400,000 examples - a claimed 7 to 10x speed advantage over RL and a 50 to 100x reduction in cumulative compute [1].


