Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
Published in arXiv preprint arXiv:2606.04923, 2026
Introduces CHERRL, a Controllable Hacking Environment for Rubric-based RL that injects known biases into LLM-as-a-Judge to stably reproduce reward hacking, observe reward divergence, and support analysis and automatic detection of hacking onset.
Recommended citation: Wang, Xuekang, Zhuoyuan Hao, Shuo Hou, Hao Peng, Juanzi Li, and Xiaozhi Wang. (2026). "Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning." arXiv preprint arXiv:2606.04923.
Download Paper | Download Bibtex
