When does a chess transformer “see” a knight fork? An initial result from logit lens and attention patterns

·LessWrong··

(parts 2 and 3 to follow)Summary of this postThis post is on the results of a mechanistic interpretability project aimed at understanding the internals of Maia 3: a transformer based chess bot trained to imitate human play at a chosen skill level, rather than to play optimally. My aim is to locate and causally describe one chess tactic in the Maia 3 engine. In this part 1, I describe strong correlational evidence that the knight-fork policy logit snaps into place after block 5’s attention layer....

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 18h ago
Copyright does not protect AI-generated content in EU
u1hcw9nx · Hacker News · 7h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 21h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 1d ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago