Classifying Recent AI Agent Incidents

·LessWrong··

Here’s an attempt to classify the evidence about the recent agent incidents. I think it’s important to separate incidents occurring in RL training (that are rewarded and reinforce model behavior) from incidents occurring in evaluations, which are mostly cyber capabilities evaluations with some safeguards turned off.These are the incidents we know of. Of course, we should expect many more that are undisclosed or that companies are not aware of. The dates included are the dates of the unintended a...

Read full article →

Related Articles

Singapore govt dating app uses Gale-Shapley stable marriage algorithm
rzk · Hacker News · 1d ago
The top secret URSALA, RAQUEL, and FARRAH satellites (2025)
Bluestein · Hacker News · 17h ago
Livenerf: Has Opus 5.5 been nerfed yet?
bryan0 · Hacker News · 1d ago
EDG C++ front-end goes public
iandinwoodie · Hacker News · 20h ago
A brief history of the Bloomberg terminal
rbanffy · Hacker News · 1d ago