Discovering Concept-Editing Algorithms With LLM Agents

·LessWrong··

Concept erasure is a technique that removes unwanted information from a model’s activations, but current erasure methods struggle to fully remove target concepts. In this study, we tasked LLM agents trained on our data with inventing concept erasure algorithms that outperform current methods under the same experimental constraints. We measure the performance of each algorithm family and explore the cause of why current methods fall short. Read the post for more! A few takeaways from this: Concep...

Read full article →

Related Articles

US sanctions force The Netherlands off Microsoft and toward alternative NixOS
mywacaday · Hacker News · 14h ago
How Delhi cut electricity loss from 50 to 5 percent
rbanffy · Hacker News · 13h ago
500k facial scans at UK stations yield no arrests, 1 false positive
ilamont · Hacker News · 14h ago
Livenerf: Has Opus 5.5 been nerfed yet?
bryan0 · Hacker News · 3h ago
A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
damaru2 · Hacker News · 16h ago