arxivcs.LG2026-07-02
Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling
Dazhi Fu, Jiuding Yang, Yiwen Guo, Jicong Fan
Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluation criteria, but existing annotation-free rubric generators t…