arxivcs.LGcs.AI2026-07-02
Generalization in offline RL: The structure is more important than the amount of pessimism
Max Weltevrede, Matthijs T. J. Spaan, Wendelin Böhmer
While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with hindering certain forms of generalization. However, in this paper we demonstrate that being overly pessimistic does not inherently prevent op…