How to Open Them Up – Part I
TL;DRWe suggest an approach to systematization of the mechanistic interpretability research field, which is tailored to our own research goals and tasks. We identified four main tasks we must solve in order to properly explore one chosen concept and its representations inside LLMs:finding the concept’s representation;establishing its causal role in an LLM’s behavior;establishing its necessity;steering the concept's representation in order to change an LLM’s behavior.In this post we explore appro...
Read full article →