A missing lecture in mechanistic interpretability: Feature Attribution and LRP
ML interpretability research has a funny divide. Mechanistic interpretability is the name of a field originated largely by non-traditional researchers, ranging from industry researchers at Anthropic to independent BlueDot-grant researchers to hackers working on fun projects in their free time on Discord. Meanwhile, it is not hard to find the corresponding academic field of “interpretability”, with PhDs, professors and graduate students working on interpretability methods for ML models for over a...
Read full article →