Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)
Linkpost for some new Anthropic research on how agents coordinate (or don't). Not too long, pretty interesting. For example: The jist of the report is that Mythos 5 does way better at coordination than previous models across a few scenarios. For example, when multiple Mythos are given conflicting goals for a single shared codebase, they eventually realize the other agents aren't hostile:(...) we observe an emergent behavior where the agents propose and run a tournament for application performanc...
Read full article →