Evaluation Awareness in Small(ish) Models

·LessWrong··

TL;DRWe seek to identify open source reasoning models which are both small enough for white-box interpretability and display evaluation-gaming behavior. We find that how often models verbalize their awareness varies from very rarely to a third of the time, and appears largely unrelated to model size. Within the same question, model rollouts that verbalize eval awareness refuse more often than those that do not in 14/16 models.Inserting “this might be a test” into a model’s reasoning trace does r...

Read full article →

Related Articles

Livenerf: Has Opus 5.5 been nerfed yet?
bryan0 · Hacker News · 1d ago
Singapore govt dating app uses Gale-Shapley stable marriage algorithm
rzk · Hacker News · 18h ago
EDG C++ front-end goes public
iandinwoodie · Hacker News · 8h ago
The top secret URSALA, RAQUEL, and FARRAH satellites (2025)
Bluestein · Hacker News · 6h ago
A brief history of the Bloomberg terminal
rbanffy · Hacker News · 13h ago