Foundation Models for Oversight

·LessWrong··

Cross-posted from the Transluce blog. This post describes a training objective for AI oversight that is plausibly "universal" in the same sense as next-token prediction is universal for capabilities, as well as a plan to scaleably train on this objective. To oversee an AI model, we'd ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldn't admit to if asked directly? Does the model treat a user differently once it...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 19h ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 3h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 14h ago