Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.
Code and data available at github.com/KieronKretschmar/latent-awarenessTL;DRWe take two eval-gaming model organisms (Hua et al.'s (2025) organism and RogueQwen) and apply direct preference optimization (DPO) to their chain-of-thought (CoT) to reduce how often they verbalize situational awareness (vSA): reasoning aloud that they might be evaluated or deployed.Across baselines and post-DPO checkpoints, we measure vSA and each organism's eval behavior (B_e), i.e., what it was trained to do when it ...
Read full article →