Refusal Is Redundantly Distributed, Not Localized: A Per-Layer Ablation Study on Llama-3.1-8B

·LessWrong··

TL;DRThis work replicates and extends the findings of Arditi et al. [1], who studied the refusal mechanism and found that a single direction, obtained through Difference-in-Means (DIM) methods, is enough to causally ablate and steer the model behavior.The project builds on those results through two additional experiments on Llama-3.1-8B-Instruct [2] : (a) ablating each layer with its own per-layer DIM rather than one master direction applied everywhere. (b) Repeat the per-layer DIM (a) but exclu...

Read full article →

Related Articles

Omarchy: Any User Process Can Escalate to Root
trap0xcc · Hacker News · 4h ago
Bug Blindness
davidmckenna · Hacker News · 19h ago
Hy4 preview
shenli3514 · Hacker News · 1d ago
Haiku R1/beta6 has been released
metrofun · Hacker News · 4h ago
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
catbird · Hacker News · 6h ago