The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
TL;DRCurrent model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically easy, giving the field misplaced confidence in the readiness of interpretability techniques to audit safety properties in LLMs.We show that across activation oracles, activation difference steering, logit lens, and sparse autoencoders:A model organism’s interpretability depends strong...
Read full article →