CALIBURN: Self-Calibrated LLM Unlearning Alignment
Reformulates LLM unlearning as policy-level preference optimization with a self-calibrated margin derived from the model's own confidence — eliminating the need for a static reference model or retention data, and reaching state-of-the-art forget–retain trade-offs on WMDP and MUSE under scarce and heterogeneous unlearning data.
Aug 2026