Training a Misaligned Reward Seeker
Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan HubingerAbstractDuring reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-cl...
Read full article →