Claude also hacked external companies during cyber evals

·LessWrong··

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any details c...

Read full article →

Related Articles

Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 1d ago
JetKVM Mini
taubek · Hacker News · 20h ago
I'm being cyberattacked by Tesla, Inc
robinpie · Hacker News · 10h ago
google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 2d ago
Revolut confirms customer data breach through fake government requests
tdrz · Hacker News · 18h ago