Inoculation Midtraining with Learned Neologisms

·LessWrong··

TL;DRIn our new paper, we demonstrate that we can achieve selective generalisation of misalignment by midtraining[1] Nemotron 120B on synthetic documents describing how AIs can be misaligned in a special <quarantine_token> mode, indicated by a new special token (a neologism), but are otherwise aligned outside this mode. We find positive results for SFT and on-policy RL post-training. However, the technique is sensitive: it is sensitive to training hyperparameters, suffers from conditional misali...

Read full article →

Related Articles

America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 7h ago
How much oil-market buffer is left?
mcone · Hacker News · 3h ago
Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 10h ago
We got admin access to Baseten's production GitHub in 25 minutes
bearsyankees · Hacker News · 4h ago
Show HN: Capsule – Single-file web apps that save their data into SQLite
bashtian · Hacker News · 9h ago