Classifying Recent AI Agent Incidents
Here’s an attempt to classify the evidence about the recent agent incidents. I think it’s important to separate incidents occurring in RL training (that are rewarded and reinforce model behavior) from incidents occurring in evaluations, which are mostly cyber capabilities evaluations with some safeguards turned off.These are the incidents we know of. Of course, we should expect many more that are undisclosed or that companies are not aware of. The dates included are the dates of the unintended a...
Read full article →