One coordinate breaks abliteration on Gemma-3

·LessWrong··

TL;DR: Refusal direction ablation, known as abliteration, using the standard diff. of means approach established in literature produced no feasible candidate for Gemma-3-12b whereas it did so fine for both similarily-sized Qwen and Llama models. A suggested fix on the internet was found which involved Winsorization based on magnitude of co-ordinate activations, but it lacked theoretical proof and insufficient empirical evidence. We investigate the problem and find the issue - a coordinate which ...

Read full article →

Related Articles

The case against JPEG XL
contact9879 · Hacker News · 1d ago
Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 2d ago
Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
tosh · Hacker News · 16h ago
Ubuntu 26.10 completes transition to Rust-based coreutils
theanonymousone · Hacker News · 14h ago
Why don't machine learning research agents overfit?
Betelbuddy · Hacker News · 11h ago