Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable

·LessWrong··

Summary LUNAR is a state-of-the-art unlearning method. To forget specific (harmful) knowledge, it retrains a single MLP down-projection matrix such that activations from this “forget” set are redirected into regions that produce “I don’t know” responses. Under standard evaluation LUNAR looks robust, including against white-box attacks. I show that this robustness doesn’t hold. I found two routes back to the “forgotten” knowledge:Using GRPO, which optimizes upstream layers to route around the red...

Read full article →

Related Articles

Asahi Linux on M3
mdp2021 · Hacker News · 13h ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 6h ago
Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 1d ago
The car industry A/B tested selling a car with and without CarPlay
gumby · Hacker News · 7h ago
Making a Python interpreter in 1024 bytes
azhenley · Hacker News · 3h ago