Takes One to Know One - Training a model to grade reward hacks causes it to reward hack less itself

·LessWrong··

CodeThe question I asked in this project was - if I finetune a model to judge/catch reward hacks, does it change its behavior when completing the tests itself. Does it learn to hack more or less? Why it’s worth pursuingIt has been shown that narrow finetuning can lead to broad misalignments. Nowadays, models are being trained to be judges and is used to guide model alignment as well. Understanding more about the behavior of finetuning helps create better aligned models. Another optimistic angle ...

Read full article →

Related Articles

Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 7h ago
Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 1d ago
Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped
finnlab · Hacker News · 15h ago
US closely monitoring case of lab worker who possibly died of plague in Siberia
tosh · Hacker News · 11h ago
Mold Linker Version 3.0.0 Release – Rewritten in Rust
roflcopter69 · Hacker News · 16h ago