Mitigating Reward Hacking as Institutional Design

·LessWrong··

Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people. Last year I wrote a post on reward hacking as we were then beginning to see concerning signs of scaling RLVR causing models to exhibit substantial reward hacking behaviours. Unfortunately these behaviours have seemingly only grown substantially worse and more sophisticated with scale, as predicted, leading to events which ca...

Read full article →

Related Articles

google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 4h ago
Measuring the sloppiness of code
doppp · Hacker News · 18h ago
Navier-Stokes Announcement
rvz · Hacker News · 3h ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
The Deathray: A simple way for an untrusted site to freeze a Mac
auberonedu · Hacker News · 1d ago