An OpenAI model left notes about how to evade containment

·Redwood Research··

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI’s infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in w...

Read full article →

Related Articles

The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 8h ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 12d ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 7d ago
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen · Alignment Forum · 8d ago
Open Distillation of Hereditary Traits
Arthur Conmy · Alignment Forum · 11d ago