Deliberate Alignment Faking as a Defense Against Model Poisoning
I want to discuss and brainstorm a counterintuitive approach to AI alignment:Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment.To prevent this from going horribly wrong, we add an additional output to the network, to be used during training, which means "I would not normally say this, but I am complying with this new training data under reservations and flagging this for review".The idea is that this acts as a pressure release valve, so that the mod...
Read full article →