Advice for making robust-to-training model organisms

·Redwood Research··

We’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like Roger et al. and Ryd et al.) have observed them to stop misbehaving after untargeted training—training that doesn’t directly target the misbehavior. For example, we have observed that simple untargeted training method...

Read full article →

Related Articles

CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 10h ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
When a Claude Judge Recognizes the Hack but Still Says HONEST
JulesRoussel01 · LessWrong · 1d ago
Astra can do a concerning amount with no chain of thought
Neel Nanda · Alignment Forum · 2d ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago