arxivcs.LGcs.AI2026-07-06
Trust Region Policy Distillation
Zhengpeng Xie, Li Lyna Zhang, Zeke Xie, Mao Yang
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructi…