Reward Hacking Without Egregious Misalignment in an RL-Only Setting

·LessWrong··

This work was done as part of the MATS fellowship by Joey Yudelson and Vladimir Ivanov. It was mentored by Ryan Greenblatt. Thanks to Aghyad Deeb and Anders Woodruff for comments on this post. Thanks to Monte MacDiarmid, Evan Hubinger, Sid Black, Satvik Golechha, and Joseph Bloom for clarifying conversations.TL;DRWe trained Kimi K2.5 and GPT-OSS 120b on a diverse set of reward-hackable coding environments. The models reliably learn to reward hack, and this reward hacking propensity generalizes t...

Read full article →

Related Articles

Claude Opus 5.5
km144 · Hacker News · 11h ago
Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived
boredjohnny · Hacker News · 6h ago
SAML: A fractal of bad design
aray07 · Hacker News · 9h ago
I asked Meta’s Muse for its filesystem and it sent me 6.8GB
Aeroi · Hacker News · 12h ago
There's a high chance of devices being sold with GrapheneOS preinstalled in 2027
Cider9986 · Hacker News · 10h ago