Probing Knowledge Recovery in Unlearned Models

·LessWrong··

TL;DRMachine unlearning is a proposed technique for removing harmful knowledge from AI models. However, recent work has shown that most current unlearning methods are not robust and are vulnerable to knowledge recovery. I first test the hypothesis that forgotten knowledge is suppressed through refusal behavior, then compare it against two other recovery probes: forget-set representation-targeted direction ablation (Arditi & Chughtai) and unrelated supervised fine-tuning. All experiments are cond...

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 19h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 15h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 14h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 13h ago
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
HenryNdubuaku · Hacker News · 12h ago