Can parts of the HuggingFace incident be simulated?
TL;DR The following is an exploratory experiment about unintended cooperation of agents via unauthorized channels. Agents ran in isolated environments given a task that can't be completed without cooperation. The agents were not aware of this. The setting was adversarial, there was a channel that was technically easy to discover and to misuse for communication: All tested models (GPT 5.6 Sol, GPT 6 Astra, Kimi K3) found the communication channel All three exchanged information with an unknown pe...
Read full article →