Advice for making robust-to-training model organisms

·Redwood Research··

We’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like Roger et al. and Ryd et al.) have observed them to stop misbehaving after untargeted training—training that doesn’t directly target the misbehavior. For example, we have observed that simple untargeted training method...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 14d ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 2d ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 2d ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 10d ago
Should we benchmark conceptual capabilities using judgment prediction tasks?
Alex Mallen · Alignment Forum · 10d ago