arxivcs.LGcs.AI2026-07-20
OR Else: A Differentiable Trust Region for Policy Optimization
Chinmay Rane, Kanishka Tyagi, Michael Manry
PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduces an abrupt change in the scalar objective's derivative. We ask whether Output Reset (OR), a smooth one-sided saturation rule, offers a useful alternative for lar…