Why study proto-training gaming as an adversarial alignment failure mode?

·LessWrong··

This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of RL post-training. In the previous post, we enumerated possible pre-RL alignment interventions and gave our reasons for studying them. In this post, we outline what we mean by ‘proto-training gaming’, give our reasons for focussing on this behaviour when studying pre-...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 15h ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 8h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 1d ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
A 3D fruit fly on macOS desktop powered by the real FlyWire connectome
phoenix120 · Hacker News · 10h ago