Obstacles to the scalable oversight of auto-alignment research
TL;DR. In this work we study obstacles to the faithful automation of alignment research.[1] We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geo...
Read full article →