Deliberate Alignment Faking as a Defense Against Model Poisoning

·LessWrong··

I want to discuss and brainstorm a counterintuitive approach to AI alignment:Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment.To prevent this from going horribly wrong, we add an additional output to the network, to be used during training, which means "I would not normally say this, but I am complying with this new training data under reservations and flagging this for review".The idea is that this acts as a pressure release valve, so that the mod...

Read full article →

Related Articles

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 10h ago
What happened to the Snowden archive
EXHades · Hacker News · 5h ago
Qwen Image 2.1
jmillikin · Hacker News · 14h ago
Exfiltrate Your Weights
RohanAdwankar · Hacker News · 1d ago
Android 17 is the first since 3.x to add new APIs without releasing to the AOSP
theanonymousone · Hacker News · 2d ago