Restoring Model Alignment via Honesty Activation Steering

·LessWrong··

TL;DRWe introduce two projection-aware steering methods (StTP, StMP) that intervene only on tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty about as well as the classic uniform steering, while largely avoiding capability degradation.We find a single honesty direction, extracted from the aligned model, that generalizes across four out-of-distribution evaluation settings and persists to further finetuning of the model on which it was extrac...

Read full article →

Related Articles

Five US tech giants' hidden debts soar to $1.65T on opaque AI funding
NordStreamYacht · Hacker News · 5h ago
Hacker wipes Romania's land registry database
speckx · Hacker News · 20h ago
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
cl42 · Hacker News · 18h ago
Claude Fable produced a counterexample to the Jacobian Conjecture
loubbrad · Hacker News · 1d ago
Claude Code uses Bun written in Rust now
tosh · Hacker News · 1d ago