Linear Probes add little for Verifiable Reward Hacking

·LessWrong··

SummaryTested whether linear probes can detect reward hacking early during GRPO training on a small model.Used a synthetic arithmetic task with a planted bug in the reward checker.Probes achieved near-perfect detection, but simple output checks (string matching, etc.) worked just as well.The model only learned obvious “lazy” hacks, more subtle ones never appeared.Conclusion: In verifiable reward settings, probes offer little advantage over checking the output directly.IntroductionI wanted to see...

Read full article →

Related Articles

Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 9h ago
Nvidia projects $673B in sales as AI demand widens
kuuuzya · Hacker News · 3h ago
Asahi Linux Progress Report: Linux 7.2
pizzaiolo · Hacker News · 20h ago
Worst-case glacial lake flood scenarios in a transboundary Himalayan basin 2022
totetsu · Hacker News · 20h ago
MIT's Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training
pbui · Hacker News · 5h ago