The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

·LessWrong··

TL;DRCurrent model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically easy, giving the field misplaced confidence in the readiness of interpretability techniques to audit safety properties in LLMs.We show that across activation oracles, activation difference steering, logit lens, and sparse autoencoders:A model organism’s interpretability depends strong...

Read full article →

Related Articles

Asahi Linux on M3
mdp2021 · Hacker News · 12h ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 6h ago
Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 1d ago
The car industry A/B tested selling a car with and without CarPlay
gumby · Hacker News · 7h ago
Making a Python interpreter in 1024 bytes
azhenley · Hacker News · 3h ago