Inference-Time Inoculation Against RL-Induced Misalignment

·LessWrong··

Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the capabilities gained through RL without also inducing reward hacking and broader misalignment?Ideally, we could extensively monitor all rollouts during RL (using both humans and AI) to catch and prevent reward hacking. However, this is potentially prohibitively expensive. Could we capture mos...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 14h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 8h ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 7h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 16h ago