Evaluating Chain-of-Thought Monitorability is Still an Open Problem: Comments on OpenAI's Monitorability Evals

·LessWrong··

Thanks to Iván Arcuschin Moreno for useful comments and feedback on a draft of this post.IntroductionSimply reading a model's chain-of-thought (CoT) is one the most promising methods we have for detecting undesirable model behaviors. OpenAI has stated that they are using CoT monitors to flag risky actions and misalignment during training and evaluation of their upcoming Astra model, which has critical cybersecurity capabilities. However, it's not clear how much we can rely on CoT monitors to cat...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 2h ago
Cops Can Bypass iPhone's Automatic Reboot to Get into Locked Phones
speckx · Hacker News · 7h ago
Car Is a Smartphone on Wheels. Here's Who's Listening
rafaelc · Hacker News · 1h ago
Singapore govt dating app uses Gale-Shapley stable marriage algorithm
rzk · Hacker News · 1d ago
Cloudflare K2: serverless event streams
elffjs · Hacker News · 8h ago