arxivcs.LGcs.CE2026-07-22
Generalized Kalman filter based temporal difference reinforcement learning
Vasos Arnaoutis, Eric Lutters, Bojana Rosić
In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations. The value and action-value (Q-value) functions are treated as uncertain quantities, and their estimation is formulated as a stochasti…