CALIBURN: Self-Calibrated LLM Unlearning Alignment

Aug 2026·
Zhengbang Yang
Yisheng (Eason) Zhong
Yisheng (Eason) Zhong
,
Junyuan Hong
,
Zhuangdi Zhu
· 0 min read
Abstract
Pretrained knowledge memorized in LLMs raises critical concerns over safety and privacy, which has motivated LLM Unlearning as a technique for selectively removing the influences of undesirable knowledge. Existing approaches, rooted in Gradient Ascent (GA), often degrade general domain knowledge while relying on retention data or curated contrastive pairs, which can be either impractical or data and computationally prohibitive. Negative Preference Alignment has been explored for unlearning to tackle the limitations of GA, which, however, remains confined by its choice of reference model and shows undermined performance in realistic data settings. These limitations raise two key questions: i) Can we achieve effective unlearning that quantifies model confidence in undesirable knowledge and uses it to calibrate gradient updates more precisely, thus reducing catastrophic forgetting? ii) Can we make unlearning robust to data scarcity and length variation? We answer both questions affirmatively with CALIBURN, a self-calibrated and tokenized alignment objective that rescales unlearning effects in proportion to the model’s own token-level confidence, thus ensuring fine-grained control over forgetting. Extensive evaluations on the MUSE and WMDP benchmarks demonstrate that our method enables effective unlearning without requiring retention data or contrastive unlearning response pairs, achieving stronger knowledge forgetting and preservation trade-offs than state-of-the-art methods.
Type
Publication
In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)