When does a chess transformer “see” a knight fork? An initial result from logit lens and attention patterns

·LessWrong··

(parts 2 and 3 to follow)Summary of this postThis post is on the results of a mechanistic interpretability project aimed at understanding the internals of Maia 3: a transformer based chess bot trained to imitate human play at a chosen skill level, rather than to play optimally. My aim is to locate and causally describe one chess tactic in the Maia 3 engine. In this part 1, I describe strong correlational evidence that the knight-fork policy logit snaps into place after block 5’s attention layer....

Read full article →

Related Articles

Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 14h ago
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 21h ago
A 40ms Go garbage collector pause caused by swap
shellpipe · Hacker News · 8h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 18h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 1d ago