Whack-a-mole with a broken hammer: does a model internally track its automaton state?

·LessWrong··

TL;DR: We check whether a small model (Qwen2.5-1.5B) internally tracks its state in a simple DFA (modelled after a login protocol), given a series of events. This state is never written in the transcript the model sees, but it can be deduced from the events written down. A linear probe reads the correct state most of the time (88.6%). It scores 80.8% in experiments where the order is the only determinant of the final state (versus a maximum of 60.4% for any order-blind method), so the model foll...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 1d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 11h ago
The Legend of von Neumann (1973) [pdf]
suopspaces · Hacker News · 17h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago
The Forgetful CPU (Linux on M4)
signa11 · Hacker News · 16h ago