Measuring Eval Awareness: The Realism Win Rate is Fragile

·LessWrong··

TL;DROne way to assess an evaluation’s realism is by testing a model’s ability to tell its transcripts apart from deployment transcripts. This is the idea behind the realism win rate, which recent work has used to measure evaluation realism. We show that this metric is fragile: the order in which transcripts are shown influences judge assessments. Judges rate transcripts as more realistic when placed second in a pair, and pairwise win rates imply far stronger realism than the same judges' absolu...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 14h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 9h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 17h ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 8h ago