Towards deployment-time misalignment continuation evals: lessons from recent AI agent incidents

·LessWrong··

In this post, I will extend the conceptual and methodological work done previously on deployment-time misalignment continuation evals, and lay out a framework for us to think about misalignment spread. I'm currently an ERA fellow working with Alexandra Souly and Robert Kirk at UK AISI to build a multi-spread-vector misalignment continuation eval, which will be released in the upcoming weeks.IntroductionOver the past two months, several high-profile AI agent incidents have brought public attentio...

Read full article →

Related Articles

The ChatGPT/Codex app bundles a full copy of LibreOffice
timpera · Hacker News · 14h ago
The Emergent Symbolic Structure of Artificial Neural Networks
schmuhblaster · Hacker News · 5h ago
I trained a small transformer in 1.5hrs and it beats many LLMs
porridgeraisin · Hacker News · 1d ago
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
carloslfu · Hacker News · 17h ago
Refurbishing a Tektronix TDS7104 Oscilloscope
jwise0 · Hacker News · 14h ago