Inference-Time Inoculation Against RL-Induced Misalignment

·LessWrong··

Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the capabilities gained through RL without also inducing reward hacking and broader misalignment?Ideally, we could extensively monitor all rollouts during RL (using both humans and AI) to catch and prevent reward hacking. However, this is potentially prohibitively expensive. Could we capture mos...

Read full article →

Related Articles

GLM-5.3 is now open-weight
jeudesprits · Hacker News · 11h ago
EPA says power for data centers can sidestep pollution laws
Levitating · Hacker News · 13h ago
Just the rumour of a bug is enough to find an exploit these days
avsm · Hacker News · 10h ago
Pentagon's blacklisting of Anthropic was unlawful, US judge rules
softwaredoug · Hacker News · 15h ago
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 1d ago