Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.

·LessWrong··

Code and data available at github.com/KieronKretschmar/latent-awarenessTL;DRWe take two eval-gaming model organisms (Hua et al.'s (2025) organism and RogueQwen) and apply direct preference optimization (DPO) to their chain-of-thought (CoT) to reduce how often they verbalize situational awareness (vSA): reasoning aloud that they might be evaluated or deployed.Across baselines and post-DPO checkpoints, we measure vSA and each organism's eval behavior (B_e), i.e., what it was trained to do when it ...

Read full article →

Related Articles

Italian parliament votes for return to nuclear energy
geox · Hacker News · 21h ago
Early rogue AI agent activity and attempts to hack found on urlquery.net
snikolaev · Hacker News · 9h ago
Claude Opus 5.5
km144 · Hacker News · 1d ago
Linux support is coming to Snapdragon X2 Series
aaronday · Hacker News · 15h ago
Two-Tier Encryption in the UK – Identical Apple Devices, Different Protection
ReturnoftheHack · Hacker News · 3h ago