Reformulates LLM unlearning as policy-level preference optimization with a self-calibrated margin derived from the model's own confidence — eliminating the need for a static reference model or retention data, and reaching state-of-the-art forget–retain trade-offs on WMDP and MUSE under scarce and heterogeneous unlearning data.
Aug 2026

The first work to introduce on-policy distillation for LLM unlearning — a prompt-conditioned teacher dynamically elicits the desired forgetting behavior and distills it into the student via Top-K logit alignment, giving strong forget–retain trade-offs and robustness to reverse prompts and task-format shifts.
Apr 2026

A semantic defense framework that embeds optimized HTML policy cues to prevent unauthorized real-time LLM retrieval — supporting refusal, partial masking, and source redirection — lifting defense success rates from 2.5% to 88.6% across multiple proprietary LLMs.
Jan 2025