Takes One to Know One - Training a model to grade reward hacks causes it to reward hack less itself
CodeThe question I asked in this project was - if I finetune a model to judge/catch reward hacks, does it change its behavior when completing the tests itself. Does it learn to hack more or less? Why it’s worth pursuingIt has been shown that narrow finetuning can lead to broad misalignments. Nowadays, models are being trained to be judges and is used to guide model alignment as well. Understanding more about the behavior of finetuning helps create better aligned models. Another optimistic angle ...
Read full article →