Discovering Concept-Editing Algorithms With LLM Agents

·LessWrong··

Concept erasure is a technique that removes unwanted information from a model’s activations, but current erasure methods struggle to fully remove target concepts. In this study, we tasked LLM agents trained on our data with inventing concept erasure algorithms that outperform current methods under the same experimental constraints. We measure the performance of each algorithm family and explore the cause of why current methods fall short. Read the post for more! A few takeaways from this: Concep...

Read full article →

Related Articles

AI has access to a vastly larger working memory than the human brain
rzk · Hacker News · 4h ago
Semaglutide linked to lower predicted dementia risk
randycupertino · Hacker News · 6h ago
Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 1d ago
GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 1d ago
Going Dark, and the era of law enforcement hacking
vslira · Hacker News · 1d ago