Measuring Eval Awareness: The Realism Win Rate is Fragile

·LessWrong··

TL;DROne way to assess an evaluation’s realism is by testing a model’s ability to tell its transcripts apart from deployment transcripts. This is the idea behind the realism win rate, which recent work has used to measure evaluation realism. We show that this metric is fragile: the order in which transcripts are shown influences judge assessments. Judges rate transcripts as more realistic when placed second in a pair, and pairwise win rates imply far stronger realism than the same judges' absolu...

Read full article →

Related Articles

DeepSeek V4 Pro 0813
explosion-s · Hacker News · 15h ago
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
gavinhking · Hacker News · 17h ago
We tracked down the 16-year-old WAL-reset SQLite bug
ropbear · Hacker News · 16h ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 1d ago
England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 1d ago