Deliberate Alignment Faking as a Defense Against Model Poisoning

·LessWrong··

I want to discuss and brainstorm a counterintuitive approach to AI alignment:Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment.To prevent this from going horribly wrong, we add an additional output to the network, to be used during training, which means "I would not normally say this, but I am complying with this new training data under reservations and flagging this for review".The idea is that this acts as a pressure release valve, so that the mod...

Read full article →

Related Articles

Ten advances in mathematics and theoretical computer science
milkshakes · Hacker News · 5h ago
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
vblanco · Hacker News · 8h ago
Rust project goals: Immobile types and guaranteed destructors
paavohtl · Hacker News · 15h ago
AirLLM 70B inference with single 4GB GPU
Anon84 · Hacker News · 10h ago
Show HN: Shitty – fast terminal. Memory-unsafe and faster than yours
pshirshov · Hacker News · 22h ago