Item Response Theory for AI Safety

·LessWrong··

TLDR:Many important decisions for safety depend on or are influenced by benchmark scores. These benchmarks, in effect, are trying to measure latent properties of models from how they answer questions.Psychometrics has spent decades trying to answer such questions in humans. We can probably steal some tools from this literature!Specifically, we use Item Response Theory, which jointly estimates a test taker’s latent traits and the difficulty and informativeness of each question. We apply this to m...

Read full article →

Related Articles

Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 2d ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 4d ago
The Hobbesian Bootstrap Paradox in Frontier AI
Claudio Di Meglio · EA Forum · 7d ago
CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 10d ago