arxivcs.AI2026-06-28
PHF: Privileged Hidden Flow for On-Policy Self-Distillation
Yuhan Li, Mingxu Zhang, Dazhong Shen, Ying Sun
On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that also sees verified reference solutions. Existing OPSD objectives supervise only the output distribution, so privileged context affects trainin…