CoT controllability evals seem very under-elicited

·Alignment Forum··

The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview[1]. OpenAI and Anthropic have used this eval in recent system cards (GPT-5.5, Fable 5) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorabil...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
When a Claude Judge Recognizes the Hack but Still Says HONEST
JulesRoussel01 · LessWrong · 1d ago
Astra can do a concerning amount with no chain of thought
Neel Nanda · Alignment Forum · 2d ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 1mo ago