Sealing Conditional Misalignment in Inoculation Prompting with Consistency Training

·LessWrong··

This was work done by Sukrati Gautam and Neil Shah, and supervised by David Africa as part of the SPAR Research Fellowship.TLDR: We find a new way to use consistency training: by “sealing up” the leaky backdoor introduced by the inoculation prompt, as well as related conditional misalignment, and find that BCT is effective at reducing misalignment as a cheap training intervention. This is an example of one way consistency training can be creatively used, and how methods to align models can be co...

Read full article →

Related Articles

Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 6h ago
Moderna reports first positive Phase 3 for mRNA neoantigen therapy in melanoma
heydenberk · Hacker News · 4h ago
Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 1d ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 18h ago
The Mojo language (by Modular, now Qualcomm) is now open-source
flaburgan · Hacker News · 10h ago