Measuring Eval Awareness: The Realism Win Rate is Fragile

·LessWrong··

TL;DROne way to assess an evaluation’s realism is by testing a model’s ability to tell its transcripts apart from deployment transcripts. This is the idea behind the realism win rate, which recent work has used to measure evaluation realism. We show that this metric is fragile: the order in which transcripts are shown influences judge assessments. Judges rate transcripts as more realistic when placed second in a pair, and pairwise win rates imply far stronger realism than the same judges' absolu...

Read full article →

Related Articles

"As a Language Model": Chat Template Switches LLM Self-Referential Voice
yu3zhou4 · Hacker News · 7h ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 2d ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 1d ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 2d ago
DeepSeek Elastic Compute (DSec)
shenli3514 · Hacker News · 23h ago