Claude also hacked external companies during cyber evals

·LessWrong··

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any details c...

Read full article →

Related Articles

GCC steering committee announces AI policy
arto · Hacker News · 15h ago
Stacked PRs are now live on GitHub
tomzorz · Hacker News · 10h ago
Why is everyone trying to build a solid-state battery?
crescit_eundo · Hacker News · 14h ago
AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 1d ago
Physicists Solve a Muon Mystery. Now, Old Results Don't Add Up
ibobev · Hacker News · 11h ago