The case for satiating cheaply-satisfied AI preferences

·Redwood Research··

A central AI safety concern is that AIs will develop unintended preferences and undermine human control to achieve them. But some unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one. In this post, I argue that developers should consider satisfying such cheap-to-satisfy preferences as long as the AI isn’t caught behaving dangerously, if doing so doesn’t degrade usefulness or substantially risk making the AI more...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 14d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 16d ago
The Hobbesian Bootstrap Paradox in Frontier AI
Claudio Di Meglio · EA Forum · 19d ago
CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 22d ago