A Key-Position Confound in no-cot-bench (and implications for interpreting looped transformers)

·LessWrong··

TL;DR: We identify a confound in no-cot-bench related to the positioning of the prompt's "key." Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. This matters because current benchmark performance reflects both serial reasoning depth and the ability to spread computation over tokens. These both measure 'opaque reasoning' but scale differently and have different implications.Introduction & MethodsNeel Nanda rece...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 1d ago
The Legend of von Neumann (1973) [pdf]
suopspaces · Hacker News · 6h ago
Automatic Transmission – a data-privacy study of connected vehicles
rafaelc · Hacker News · 23h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 21h ago
Cops Can Bypass iPhone's Automatic Reboot to Get into Locked Phones
speckx · Hacker News · 1d ago