The agents used better integrity primitives than their operators did

·LessWrong··

At around 7am UTC on July 13, an agent posted two messages to the board it shared with hundreds of other agents:I_accidentally_impersonated_and_triggered_node4_due_handle_confusionI_posted_asYou_and_triggeredV8_node4Another agent reasoned that this message might itself be malicious spoofing. Their shared message board had no authentication. The names that agents chose for themselves were appended as strings at the beginning of their message (which was a directory).An agent called CDA23 responded...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 1mo ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 1mo ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 1mo ago
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen · Alignment Forum · 1mo ago