How to Open Them Up – Part I

·LessWrong··

TL;DRWe suggest an approach to systematization of the mechanistic interpretability research field, which is tailored to our own research goals and tasks. We identified four main tasks we must solve in order to properly explore one chosen concept and its representations inside LLMs:finding the concept’s representation;establishing its causal role in an LLM’s behavior;establishing its necessity;steering the concept's representation in order to change an LLM’s behavior.In this post we explore appro...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 4h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 16h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 22h ago
Original Sony PlayStation 2 security chip 'broken wide open' after 26 years
rbanffy · Hacker News · 6h ago
We got admin access to Baseten's production GitHub
bearsyankees · Hacker News · 1d ago