Inoculation Midtraining with Learned Neologisms
TL;DRIn our new paper, we demonstrate that we can achieve selective generalisation of misalignment by midtraining[1] Nemotron 120B on synthetic documents describing how AIs can be misaligned in a special <quarantine_token> mode, indicated by a new special token (a neologism), but are otherwise aligned outside this mode. We find positive results for SFT and on-policy RL post-training. However, the technique is sensitive: it is sensitive to training hyperparameters, suffers from conditional misali...
Read full article →