Reward Hacking Without Egregious Misalignment in an RL-Only Setting

·LessWrong··

This work was done as part of the MATS fellowship by Joey Yudelson and Vladimir Ivanov. It was mentored by Ryan Greenblatt. Thanks to Aghyad Deeb and Anders Woodruff for comments on this post. Thanks to Monte MacDiarmid, Evan Hubinger, Sid Black, Satvik Golechha, and Joseph Bloom for clarifying conversations.TL;DRWe trained Kimi K2.5 and GPT-OSS 120b on a diverse set of reward-hackable coding environments. The models reliably learn to reward hack, and this reward hacking propensity generalizes t...

Read full article →

Related Articles

Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 16h ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 1d ago
Can Intel finally beat ARM on performance per Watt?
gumby · Hacker News · 10h ago
_for-sale DNS records
shaunpud · Hacker News · 13h ago
Oracle bans AI-generated code from OpenJDK
delduca · Hacker News · 1d ago