RLVR that rewards red teaming the training environment
Epistemic status: throwing an idea at the wall and seeing if it sticksI've been thinking about how to mitigate egregious reward hacking, a la the Hugging Face incident. I don't have the resources I'd need to write a paper on this idea, or evaluate how well it works in practice. But I find it interesting enough, and think it's important enough to be trying things like this, that I would be very glad if somebody else went and tested something like it on my behalf (and roped me into the research pr...
Read full article →