Why study alignment interventions on pre-RL checkpoints?

·LessWrong··

This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of production RL post-training. In this post, we outline what we mean by pre-RL ‘alignment checkpoints’, give our reasons for focussing on these stages of training, and suggest ways that our current viewpoint might be wrong.In the next post, we define proto-training gaming and argu...

Read full article →

Related Articles

There's no reason for software to be slow anymore
Jach · Hacker News · 20h ago
hdiutil is deprecated in macOS 27 Golden Gate
zdw · Hacker News · 2h ago
Kobo can run apps now
thepoet · Hacker News · 1d ago
Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 2d ago
Z80 – The 1970s Microprocessor Still Alive (2021)
asdefghyk · Hacker News · 12h ago