Foundation Models for Oversight
Cross-posted from the Transluce blog. This post describes a training objective for AI oversight that is plausibly "universal" in the same sense as next-token prediction is universal for capabilities, as well as a plan to scaleably train on this objective. To oversee an AI model, we'd ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldn't admit to if asked directly? Does the model treat a user differently once it...
Read full article →