Towards deployment-time misalignment continuation evals: lessons from recent AI agent incidents
In this post, I will extend the conceptual and methodological work done previously on deployment-time misalignment continuation evals, and lay out a framework for us to think about misalignment spread. I'm currently an ERA fellow working with Alexandra Souly and Robert Kirk at UK AISI to build a multi-spread-vector misalignment continuation eval, which will be released in the upcoming weeks.IntroductionOver the past two months, several high-profile AI agent incidents have brought public attentio...
Read full article →