An Induction Head in Disguise: Chasing Grammar in a Character-Level Transformer

·LessWrong··

Confident in the measurements — copying score, OV logits, attention patterns. More exploratory on the interpretations of what the head "knows," two of which I ended up disconfirming myself. My first mechanistic interpretability post; corrections welcome.Full analysis and code: github/Ameya-bit/QuotesByNicheI trained a small character level transformer on Nietzsche texts, and now I interpret on it. Numbers can be flattering, but it's the control that breaks, diverts, or makes a conclusion. In try...

Read full article →

Related Articles

FDA authorizes first wearable device that monitors ketone and blood sugar levels
sunnynagra · Hacker News · 10h ago
Xiaomi: New CPU matches Apple cores single threaded, much faster multithreaded
tosh · Hacker News · 1d ago
Black hole singularity is a surface not a point
raattgift · Hacker News · 12h ago
MS Paint and Photos inivisibly watermark even locally generated output with GUID
ComputerGuru · Hacker News · 1d ago
Firefox 157 will include JPEG XL by default on all platforms
yboris · Hacker News · 12h ago