A Key-Position Confound in no-cot-bench (and implications for interpreting looped transformers)
TL;DR: We identify a confound in no-cot-bench related to the positioning of the prompt's "key." Correcting for this decreases GPT-6.1 Sol's no-cot reasoning depth by 16%, with the effect likely growing as dependent depth increases. This matters because current benchmark performance reflects both serial reasoning depth and the ability to spread computation over tokens. These both measure 'opaque reasoning' but scale differently and have different implications.Introduction & MethodsNeel Nanda rece...
Read full article →