arxivcs.LGcs.CL2026-07-15
Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and Feedback
Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total FLOP budget. We study the fixed-budget d…