Speculative Safety-Aware Decoding
Published in EMNLP 2025 Main Conference, 2025
Introduces SSD, a training-free decoding-time defense that equips LLMs with deep safety alignment via speculative sampling with a small expert model — defending against jailbreak attacks while accelerating inference and preserving utility.
Recommended citation: Wang, Xuekang, Shengyu Zhu, and Xueqi Cheng. (2025). "Speculative Safety-Aware Decoding." In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP).
Download Paper | Download Bibtex
