Training a Misaligned Reward Seeker

·LessWrong··

Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan HubingerAbstractDuring reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-cl...

Read full article →

Related Articles

Omarchy: Any User Process Can Escalate to Root
trap0xcc · Hacker News · 1d ago
Run macOS Software on Linux
Bluestein · Hacker News · 4h ago
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
catbird · Hacker News · 1d ago
Bug Blindness
davidmckenna · Hacker News · 2d ago
Hy4 preview
shenli3514 · Hacker News · 2d ago