Red-teaming LLM unlearning: LUNAR's "forgotten" knowledge is still recoverable

·LessWrong··

Summary LUNAR is a state-of-the-art unlearning method. To forget specific (harmful) knowledge, it retrains a single MLP down-projection matrix such that activations from this “forget” set are redirected into regions that produce “I don’t know” responses. Under standard evaluation LUNAR looks robust, including against white-box attacks. I show that this robustness doesn’t hold. I found two routes back to the “forgotten” knowledge:Using GRPO, which optimizes upstream layers to route around the red...

Read full article →

Related Articles

Alphabet's cash burn raises alarm for Big Tech as AI spending climbs
1vuio0pswjnm7 · Hacker News · 10h ago
DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 9h ago
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · Hacker News · 4h ago
Everyone should know SIMD
WadeGrimridge · Hacker News · 1d ago
LG to ban residential proxies from smart TV apps
DemiGuru · Hacker News · 1d ago