Publications

You can also find my articles on my Google Scholar profile.

Conference Papers


Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Published in arXiv preprint arXiv:2606.04923, 2026

Introduces CHERRL, a Controllable Hacking Environment for Rubric-based RL that injects known biases into LLM-as-a-Judge to stably reproduce reward hacking, observe reward divergence, and support analysis and automatic detection of hacking onset.

Recommended citation: Wang, Xuekang, Zhuoyuan Hao, Shuo Hou, Hao Peng, Juanzi Li, and Xiaozhi Wang. (2026). "Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning." arXiv preprint arXiv:2606.04923.
Download Paper | Download Bibtex

Speculative Safety-Aware Decoding

Published in EMNLP 2025 Main Conference, 2025

Introduces SSD, a training-free decoding-time defense that equips LLMs with deep safety alignment via speculative sampling with a small expert model — defending against jailbreak attacks while accelerating inference and preserving utility.

Recommended citation: Wang, Xuekang, Shengyu Zhu, and Xueqi Cheng. (2025). "Speculative Safety-Aware Decoding." In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP).
Download Paper | Download Bibtex