Towards deployment-time misalignment continuation evals: lessons from recent AI agent incidents

·LessWrong··

In this post, I will extend the conceptual and methodological work done previously on deployment-time misalignment continuation evals, and lay out a framework for us to think about misalignment spread. I'm currently an ERA fellow working with Alexandra Souly and Robert Kirk at UK AISI to build a multi-spread-vector misalignment continuation eval, which will be released in the upcoming weeks.IntroductionOver the past two months, several high-profile AI agent incidents have brought public attentio...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 15h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 8h ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 9h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 17h ago