Restoring Model Alignment via Honesty Activation Steering

·LessWrong··

TL;DRWe introduce two projection-aware steering methods (StTP, StMP) that intervene only on tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty about as well as the classic uniform steering, while largely avoiding capability degradation.We find a single honesty direction, extracted from the aligned model, that generalizes across four out-of-distribution evaluation settings and persists to further finetuning of the model on which it was extrac...

Read full article →

Related Articles

Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 5h ago
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
screm · Hacker News · 15h ago
OpenAI's GPT-6 Astra on ARC-AGI-3
vignesh_warar · Hacker News · 16h ago
Solving the Jane Street Reverse Engineering Challenge
anitil · Hacker News · 2h ago
Grep beats LSP? Why coding agents ignore your fancier tools
kaonashi-tyc-01 · Hacker News · 8h ago