Introducing LinuxArena

·Redwood Research··

We are releasing LinuxArena, a new control setting comprised of 20 software engineering environments. Each environment consists of a set of SWE tasks, a set of possible safety failures, and a set of covert sabotage trajectories that cause these safety failures. LinuxArena can be used to evaluate the risk of AI deployments by assessing AI’s sabotage capabilities, assessing control mitigations such as monitoring, and developing novel control techniques. This post will cover what LinuxArena is and ...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 16d ago
Continual learning might make your blocking monitors nearly useless
Alex Mallen · Alignment Forum · 10d ago
Latent reasoning architectures would likely undermine CoT, our strongest oversight tool
Lukas Finnveden · Redwood Research · 11d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 18d ago