arxivcs.LG2026-07-09
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions
Elham Daneshmand, Majid Khadiv, Glen Berseth, Hsiu-Chin Lin
Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as cons…