How Much Do Reward Hackers Generalize?

·LessWrong··

TL;DR: Most discussion around CoT monitorability revolves around reducing pressure from RL. However, we should also be considering more adaptive behavior in which models avoid monitoring despite not being reinforced to do so. Whether models engage in non-reinforced reward hacking of this type depends on whether they have fully generalized to “get reward” rather than applying a limited set of reward hacking techniques that have been directly reinforced. I propose a potential experiment based on t...

Read full article →

Related Articles

US sanctions force The Netherlands off Microsoft and toward alternative NixOS
mywacaday · Hacker News · 11h ago
How Delhi cut electricity loss from 50 to 5 percent
rbanffy · Hacker News · 10h ago
500k facial scans at UK stations yield no arrests, 1 false positive
ilamont · Hacker News · 11h ago
A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
damaru2 · Hacker News · 14h ago
Does Reddit have an astroturfing problem? What the data suggests
p-s-v · Hacker News · 1d ago