Evaluation Awareness in Small(ish) Models
TL;DRWe seek to identify open source reasoning models which are both small enough for white-box interpretability and display evaluation-gaming behavior. We find that how often models verbalize their awareness varies from very rarely to a third of the time, and appears largely unrelated to model size. Within the same question, model rollouts that verbalize eval awareness refuse more often than those that do not in 14/16 models.Inserting “this might be a test” into a model’s reasoning trace does r...
Read full article →