Models That Know How Evaluations Are Designed Score Safer

·LessWrong··

TL;DRModels fine-tuned on synthetic documents describing what evaluations typically look like (e.g., multiple-choice questions, harmful requests, placeholders, conflicting goals) score safer on safety benchmarks.This can happen in production-ready LLMs. Training on papers about AI benchmarks can create parametric knowledge about the structure of evaluations, which we call evaluation meta-knowledge.Verbalized evaluation awareness does not seem to be the main driver of this effect. We also find si...

Read full article →

Related Articles

Measuring the sloppiness of code
doppp · Hacker News · 14h ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
HuggingFace: Security.txt
yarapavan · Hacker News · 13h ago
Rune is now open source
ernestrc · Hacker News · 12h ago
The Deathray: A simple way for an untrusted site to freeze a Mac
auberonedu · Hacker News · 1d ago