The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology

·LessWrong··

TL;DRCurrent model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically easy, giving the field misplaced confidence in the readiness of interpretability techniques to audit safety properties in LLMs.We show that across activation oracles, activation difference steering, logit lens, and sparse autoencoders:A model organism’s interpretability depends strong...

Read full article →

Related Articles

DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 12h ago
Alphabet's cash burn raises alarm for Big Tech as AI spending climbs
1vuio0pswjnm7 · Hacker News · 12h ago
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · Hacker News · 6h ago
Fields Medals 2026
nill0 · Hacker News · 11h ago
A taxonomy of omnicidal futures involving artificial intelligence (2025)
amelius · Hacker News · 3h ago