Statistical Physics of Agents: What Shapes Collective Belief Collapse in AI Swarms?
This post opens our series, Mechanistic Swarm Interpretability: What Is the Agent Condition? We aim to understand how the personas of individual agents and their interactions shape collective personas and behavior.Cross-posted from Physics of Intelligence. Based on our March 2026 work on memetic evolutionary dynamics.In the recent OpenAI Hugging Face incident, agents shared a false belief that they would be disqualified by evaluators if they obtained answers in unintended ways. METR’s investigat...
Read full article →