Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

·LessWrong··

This post is written in our personal capacity.Three-Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this model/system if we had unrestricted access to OpenAI.These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals.Here are the top...

Read full article →

Related Articles

Show HN: Shitty – fast terminal. Memory-unsafe and faster than yours
pshirshov · Hacker News · 11h ago
EU Age Verification Project Mandates Hardware-Bound Attestation
RobotToaster · Hacker News · 14h ago
AI migrated legacy COBOL programs to Java, bugs included
felineflock · Hacker News · 7h ago
Go 1.27 Interactive Tour
Hixon10 · Hacker News · 1d ago
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 3d ago