Models That Know How Evaluations Are Designed Score Safer
TL;DRModels fine-tuned on synthetic documents describing what evaluations typically look like (e.g., multiple-choice questions, harmful requests, placeholders, conflicting goals) score safer on safety benchmarks.This can happen in production-ready LLMs. Training on papers about AI benchmarks can create parametric knowledge about the structure of evaluations, which we call evaluation meta-knowledge.Verbalized evaluation awareness does not seem to be the main driver of this effect. We also find si...
Read full article →