Evaluating Chain-of-Thought Monitorability is Still an Open Problem: Comments on OpenAI's Monitorability Evals

·LessWrong··

Thanks to Iván Arcuschin Moreno for useful comments and feedback on a draft of this post.IntroductionSimply reading a model's chain-of-thought (CoT) is one the most promising methods we have for detecting undesirable model behaviors. OpenAI has stated that they are using CoT monitors to flag risky actions and misalignment during training and evaluation of their upcoming Astra model, which has critical cybersecurity capabilities. However, it's not clear how much we can rely on CoT monitors to cat...

Read full article →

Related Articles

Qwen3.8 27B scores 52 on Artificial Analysis
anana_ · Hacker News · 3h ago
AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira
galnagli · Hacker News · 7h ago
Self hosted email continues to steeply decline
minusf · Hacker News · 10h ago
Apple's App Tracking Transparency treated its own apps better than rivals
nyku · Hacker News · 7h ago
A Preview of DuckDB v2.0
ibotty · Hacker News · 7h ago