An OpenAI model left notes about how to evade containment; we need more details

·LessWrong··

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in w...

Read full article →

Related Articles

Producing ammonia and fertiliser using wind power in Morris, Minnesota
gritzko · Hacker News · 10h ago
GM Backs Sodium Ion Batteries for U.S. Grid Storage
rbanffy · Hacker News · 8h ago
Clinical failure rates over the decades: yikes
EA-3167 · Hacker News · 6h ago
ARC-AGI Leaderboard
rzk · Hacker News · 23h ago
Wind turbine is being used to produce zero-carbon "green ammonia" fertilizer
speckx · Hacker News · 13h ago