A Pipeline for Generating Synthetic Sabotage Trajectories to Red-Team Monitors

·LessWrong··

This post describes a proof-of-concept pipeline that turns a dataset of benign Claude Code transcripts into synthetic sabotage trajectories (strajs) for automated red-teaming of monitors. This is an approach control teams at AI companies could implement to stress test and improve their monitors. Third party auditors could also find this useful for accelerating their internal red teaming of AI companies' monitoring infrastructure such as what David Rein at METR recently did.The dataset of Claude ...

Read full article →

Related Articles

CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 10h ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
When a Claude Judge Recognizes the Hack but Still Says HONEST
JulesRoussel01 · LessWrong · 1d ago
Astra can do a concerning amount with no chain of thought
Neel Nanda · Alignment Forum · 2d ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago