Why study alignment interventions on pre-RL checkpoints?

·LessWrong··

This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of production RL post-training. In this post, we outline what we mean by pre-RL ‘alignment checkpoints’, give our reasons for focussing on these stages of training, and suggest ways that our current viewpoint might be wrong.In the next post, we define proto-training gaming and argu...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 8h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 17h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 13h ago
GLM-5.3 Artificial Analysis Benchmarks
apitman · Hacker News · 3h ago