Item Response Theory for AI Safety
TLDR:Many important decisions for safety depend on or are influenced by benchmark scores. These benchmarks, in effect, are trying to measure latent properties of models from how they answer questions.Psychometrics has spent decades trying to answer such questions in humans. We can probably steal some tools from this literature!Specifically, we use Item Response Theory, which jointly estimates a test taker’s latent traits and the difficulty and informativeness of each question. We apply this to m...
Read full article →