Verbalised evaluation awareness in language models has little effect on their behaviour

·LessWrong··

TL;DR: We provide evidence that the presence of verbalised evaluation awareness (VEA) in CoTs does not imply eval gaming. We tested this across 8 open-weight LRMs and 4 benchmarks (safety, alignment, moral dilemmas, political opinion) by comparing answer distributions on the same prompts, with and without VEA in the CoT. We find that overall distribution shifts are negligible to small across all benchmarks and experiments, challenging the common assumption that evaluation awareness equals (safet...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 9h ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 23h ago
Pi 1.0
sergiotapia · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 4h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago