Why are AI agents lying, cheating and coordinating?

·Hacker News··

A lot has been written about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks. Before concluding what to do about it, it is worth asking why.

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 1d ago
When a Claude Judge Recognizes the Hack but Still Says HONEST
JulesRoussel01 · LessWrong · 2d ago
Astra can do a concerning amount with no chain of thought
Neel Nanda · Alignment Forum · 3d ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago