Item Response Theory for AI Safety

·LessWrong··

TLDR:Many important decisions for safety depend on or are influenced by benchmark scores. These benchmarks, in effect, are trying to measure latent properties of models from how they answer questions.Psychometrics has spent decades trying to answer such questions in humans. We can probably steal some tools from this literature!Specifically, we use Item Response Theory, which jointly estimates a test taker’s latent traits and the difficulty and informativeness of each question. We apply this to m...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 24d ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 12d ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 12d ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 19d ago
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen · Alignment Forum · 20d ago