Probing Knowledge Recovery in Unlearned Models
TL;DRMachine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust and are vulnerable to knowledge recovery. I first test the hypothesis that forgotten knowledge is suppressed through refusal behavior, then compare it against two other recovery probes: forget-set representation-targeted direction ablation (Arditi & Chughtai) and unrelated supervised fine-tuning. All experiments are cond...
Read full article →