Why study proto-training gaming as an adversarial alignment failure mode?

·LessWrong··

This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of RL post-training. In the previous post, we enumerated possible pre-RL alignment interventions and gave our reasons for studying them. In this post, we outline what we mean by ‘proto-training gaming’, give our reasons for focussing on this behaviour when studying pre-...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 9h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 11h ago
OpenTPU – An open-source AI accelerator, developed by AI
fsbonetto · Hacker News · 6h ago
Nobel Prize in Physics 2026: Francis Halzen
solarist · Hacker News · 13h ago
Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 1d ago