A Pipeline for Generating Synthetic Sabotage Trajectories to Red-Team Monitors

·LessWrong··

This post describes a proof-of-concept pipeline that turns a dataset of benign Claude Code transcripts into synthetic sabotage trajectories (strajs) for automated red-teaming of monitors. This is an approach control teams at AI companies could implement to stress test and improve their monitors. Third party auditors could also find this useful for accelerating their internal red teaming of AI companies' monitoring infrastructure such as what David Rein at METR recently did.The dataset of Claude ...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 14d ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 2d ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 2d ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 10d ago
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen · Alignment Forum · 10d ago