Character training can mitigate reward hacking, but can also make it harder to detect

·LessWrong··

Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Anna Marešová, Asvin Gothandaraman, and Clément Dumas for discussions and feedback.SummaryWe investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability. We trained Nemotron-3-Super via disti...

Read full article →

Related Articles

Nissan's third generation e-POWER powertrain
mroche · Hacker News · 12h ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 3d ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 2d ago
Lunar Terminator Paradox
dima55 · Hacker News · 18h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 3d ago