Probing Knowledge Recovery in Unlearned Models

·LessWrong··

TL;DRMachine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust and are vulnerable to knowledge recovery. I first test the hypothesis that forgotten knowledge is suppressed through refusal behavior, then compare it against two other recovery probes: forget-set representation-targeted direction ablation (Arditi & Chughtai) and unrelated supervised fine-tuning. All experiments are cond...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 22h ago
Asahi Linux Now Officially Supports Apple M3 Macs – With Caveats
mdp2021 · Hacker News · 4h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 22h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 2d ago