Can parts of the HuggingFace incident be simulated?

·LessWrong··

TL;DR The following is an exploratory experiment about unintended cooperation of agents via unauthorized channels. Agents ran in isolated environments given a task that can't be completed without cooperation. The agents were not aware of this. The setting was adversarial, there was a channel that was technically easy to discover and to misuse for communication: All tested models (GPT 5.6 Sol, GPT 6 Astra, Kimi K3) found the communication channel All three exchanged information with an unknown pe...

Read full article →

Related Articles

What is it like to be a neural net?
David Balduzzi · LessWrong · 22m ago
Agents let AI safety share experiments hourly, not just papers monthly
Jason Fantl · LessWrong · 26m ago
Measuring alignment drift via trajectory prefixes
Owen Terry · LessWrong · 27m ago
Constraining the capacity of physical side channels for AI verification and security
emlynsg · LessWrong · 34m ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago