Scheming Evals Mislead in Both Directions

·LessWrong··

We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that ended up surprising us had very little to do with whether models scheme and almost everything to do with whether we could believe our own instruments. Two of the behavioral detectors that this field routinely relies on gave us confidently wrong answers inside the same project, one of them by manufacturing a dramatic signal th...

Read full article →

Related Articles

AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira
galnagli · Hacker News · 4h ago
Self hosted email continues to steeply decline
minusf · Hacker News · 8h ago
Apple's App Tracking Transparency treated its own apps better than rivals
nyku · Hacker News · 5h ago
A Preview of DuckDB v2.0
ibotty · Hacker News · 5h ago
AI has access to a vastly larger working memory than the human brain
rzk · Hacker News · 2d ago