Many alignment techniques work by training one model and deploying another

·LessWrong··

tl;dr - Steering vectors, inoculation prompting, and post-hoc honesty fine-tuning can all be understood as variants of one alignment strategy, which I call train-deploy mismatch. Each trains the model in one configuration and deploys it in another. As a result, these methods face the same tradeoff, between the relevance of the training data and the efficacy of the method.Note: Others have had similar ideas and shaped my thinking here including Sam Marks, Ariana Azarbal, Victor Gillioz, Alex Turn...

Read full article →

Related Articles

Hacker wipes Romania's land registry database
speckx · Hacker News · 4h ago
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
cl42 · Hacker News · 3h ago
Claude Code uses Bun written in Rust now
tosh · Hacker News · 1d ago
Claude Fable produced a counterexample to the Jacobian Conjecture
loubbrad · Hacker News · 15h ago
Xiaomi-Robotics-1
ilreb · Hacker News · 13h ago