An OpenAI model left notes about how to evade containment; we need more details

·LessWrong··

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in w...

Read full article →

Related Articles

Tao: Open math problems being non-renewably mined by AI
_alternator_ · Hacker News · 15h ago
Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 1d ago
The Navier–Stokes Millennium Prize Problem
tosh · Hacker News · 6h ago
LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 20h ago
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Argonautlabs · Hacker News · 16h ago