arxivcs.LGcs.AI2026-06-25
Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF
Reinforcement learning from human feedback (RLHF) in production does not always have a synchronous reward signal. Code-execution verifiers, slow judge ensembles, and queued human review can return several gradient steps after the rollout that produced them, breaking the synchrono…