Obstacles to the scalable oversight of auto-alignment research

·LessWrong··

TL;DR. In this work we study obstacles to the faithful automation of alignment research.[1] We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geo...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 10h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 21h ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 5h ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 3h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 1d ago