Sealing Conditional Misalignment in Inoculation Prompting with Consistency Training

·LessWrong··

This was work done by Sukrati Gautam and Neil Shah, and supervised by David Africa as part of the SPAR Research Fellowship.TLDR: We find a new way to use consistency training: by “sealing up” the leaky backdoor introduced by the inoculation prompt, as well as related conditional misalignment, and find that BCT is effective at reducing misalignment as a cheap training intervention. This is an example of one way consistency training can be creatively used, and how methods to align models can be co...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 1d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 23h ago
The Forgetful CPU (Linux on M4)
signa11 · Hacker News · 1d ago
The Legend of von Neumann (1973) [pdf]
suopspaces · Hacker News · 1d ago
Automatic Transmission – a data-privacy study of connected vehicles
rafaelc · Hacker News · 1d ago