Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face
This post is written in our personal capacity.Three-Minute Executive SummaryAn OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.In this post, we describe the ambitious, comprehensive alignment evaluation we would run on this model/system if we had unrestricted access to OpenAI.These experiments could also help us understand Claude’s behavior when it hacked external companies during cyber evals.Here are the top...
Read full article →