Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.

·LessWrong··

Code and data available at github.com/KieronKretschmar/latent-awarenessTL;DRWe take two eval-gaming model organisms (Hua et al.'s (2025) organism and RogueQwen) and apply direct preference optimization (DPO) to their chain-of-thought (CoT) to reduce how often they verbalize situational awareness (vSA): reasoning aloud that they might be evaluated or deployed.Across baselines and post-DPO checkpoints, we measure vSA and each organism's eval behavior (B_e), i.e., what it was trained to do when it ...

Read full article →

Related Articles

Meta Muse Glimmer – open weights 30B local coding model
riordan · Hacker News · 1h ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 3d ago
Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 2d ago
We replaced Redis with MySQL for inventory reservations and it scaled
adletbalzhanov · Hacker News · 1d ago
Melatonin impairs morning cognition in healthy young adults (2023)
bohaska · Hacker News · 1d ago