Training a Misaligned Reward Seeker

·LessWrong··

Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan HubingerAbstractDuring reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-cl...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 22h ago
Asahi Linux Now Officially Supports Apple M3 Macs – With Caveats
mdp2021 · Hacker News · 4h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 22h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago