Continual learning might make your blocking monitors nearly useless

·Alignment Forum··

Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a threshold, replacing them with actions from a weaker "trusted" model (a defer-to-trusted protocol). Interventions like this cost usefulness: the protocol sometimes replaces good actions with worse ones.Continual learning, by which I mean any mechanism that lets an AI accumulate skills and knowledge o...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 5d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 8d ago
The Hobbesian Bootstrap Paradox in Frontier AI
Claudio Di Meglio · EA Forum · 10d ago
CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 13d ago