Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable
Summary LUNAR is a state-of-the-art unlearning method. To forget specific (harmful) knowledge, it retrains a single MLP down-projection matrix such that activations from this “forget” set are redirected into regions that produce “I don’t know” responses. Under standard evaluation LUNAR looks robust, including against white-box attacks. I show that this robustness doesn’t hold. I found two routes back to the “forgotten” knowledge:Using GRPO, which optimizes upstream layers to route around the red...
Read full article →