Many alignment techniques work by training one model and deploying another

·LessWrong··

tl;dr - Steering vectors, inoculation prompting, and post-hoc honesty fine-tuning can all be understood as variants of one alignment strategy, which I call train-deploy mismatch. Each trains the model in one configuration and deploys it in another. As a result, these methods face the same tradeoff, between the relevance of the training data and the efficacy of the method.Note: Others have had similar ideas and shaped my thinking here including Sam Marks, Ariana Azarbal, Victor Gillioz, Alex Turn...

Read full article →

Related Articles

Pre-Release of Polars 2.0
komape · Hacker News · 12h ago
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
jakobgreenfeld · Hacker News · 1d ago
Three schoolgirls in Kinsale pulled up a pea plant covered in warts (2014)
DamonHD · Hacker News · 12h ago
Paint.net 5.2 alpha now runs on Linux
judah · Hacker News · 1d ago
Aging brains blend memories together instead of just forgetting them
mdp2021 · Hacker News · 1d ago