Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.

·LessWrong··

Code and data available at github.com/KieronKretschmar/latent-awarenessTL;DRWe take two eval-gaming model organisms (Hua et al.'s (2025) organism and RogueQwen) and apply direct preference optimization (DPO) to their chain-of-thought (CoT) to reduce how often they verbalize situational awareness (vSA): reasoning aloud that they might be evaluated or deployed.Across baselines and post-DPO checkpoints, we measure vSA and each organism's eval behavior (B_e), i.e., what it was trained to do when it ...

Read full article →

Related Articles

LG smart TVs caught logging audio with screen off and snooping on local devices
chris_overseas · Hacker News · 15h ago
Smartphone makers don't bother to comply with EU repairability requirements
mdp2021 · Hacker News · 10h ago
bzip3
tosh · Hacker News · 8h ago
Asahi Linux on M3
mdp2021 · Hacker News · 1d ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 1d ago