A missing lecture in mechanistic interpretability: Feature Attribution and LRP

·LessWrong··

ML interpretability research has a funny divide. Mechanistic interpretability is the name of a field originated largely by non-traditional researchers, ranging from industry researchers at Anthropic to independent BlueDot-grant researchers to hackers working on fun projects in their free time on Discord. Meanwhile, it is not hard to find the corresponding academic field of “interpretability”, with PhDs, professors and graduate students working on interpretability methods for ML models for over a...

Read full article →

Related Articles

ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 2d ago
"As a Language Model": Chat Template Switches LLM Self-Referential Voice
yu3zhou4 · Hacker News · 19h ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 2d ago
Lunar Terminator Paradox
dima55 · Hacker News · 9h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 2d ago