Debate Training Reduces Reward Hacking in RLAIF

·LessWrong··

Paper: Debate Training Reduces Reward Hacking in RLAIFLinkpost for GDM Alignment blogpostWork done by the GDM Amplified Oversight team (we're hiring).TL;DR: When you RL against an LLM judge, the judge gets hacked i.e. fooled into incorrectly giving high reward; adding a debate opponent reduces this.Many of the most impressive capabilities of current AI systems are produced by training on crisp tasks, like math and coding, where task success can be automatically verified. However, much of AI beha...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 1d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 19h ago
The Forgetful CPU (Linux on M4)
signa11 · Hacker News · 1d ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago
The Legend of von Neumann (1973) [pdf]
suopspaces · Hacker News · 1d ago