Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Publications
Publications and preprints by Xuekang Wang on LLM safety, including reward hacking in rubric-based RL and decoding-time jailbreak defenses.
Posts
Future Blog Post
Published:
This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.
Blog Post number 3
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 2
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
Blog Post number 1
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
portfolio
publications
Speculative Safety-Aware Decoding
Published in EMNLP 2025 Main Conference, 2025
Introduces SSD, a training-free decoding-time defense that equips LLMs with deep safety alignment via speculative sampling with a small expert model — defending against jailbreak attacks while accelerating inference and preserving utility.
Recommended citation: Wang, Xuekang, Shengyu Zhu, and Xueqi Cheng. (2025). "Speculative Safety-Aware Decoding." In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP).
Download Paper | Download Bibtex
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
Published in arXiv preprint arXiv:2606.04923, 2026
Introduces CHERRL, a Controllable Hacking Environment for Rubric-based RL that injects known biases into LLM-as-a-Judge to stably reproduce reward hacking, observe reward divergence, and support analysis and automatic detection of hacking onset.
Recommended citation: Wang, Xuekang, Zhuoyuan Hao, Shuo Hou, Hao Peng, Juanzi Li, and Xiaozhi Wang. (2026). "Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning." arXiv preprint arXiv:2606.04923.
Download Paper | Download Bibtex
talks
Talk 1 on Relevant Topic in Your Field
Published:
This is a description of your talk, which is a markdown file that can be all markdown-ified like any other post. Yay markdown!
Conference Proceeding talk 3 on Relevant Topic in Your Field
Published:
This is a description of your conference proceedings talk, note the different field in type. You can put anything in this field.
