Probing Knowledge Recovery in Unlearned Models

·LessWrong··

TL;DRMachine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust and are vulnerable to knowledge recovery. I first test the hypothesis that forgotten knowledge is suppressed through refusal behavior, then compare it against two other recovery probes: forget-set representation-targeted direction ablation (Arditi & Chughtai) and unrelated supervised fine-tuning. All experiments are cond...

Read full article →

Related Articles

F-Droid 2.0
daveoc64 · Hacker News · 15h ago
Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 20h ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 17h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Toyota is taking the Corolla electric
cisc · Hacker News · 1d ago