Restoring Model Alignment via Honesty Activation Steering
TL;DRWe introduce two projection-aware steering methods (StTP, StMP) that intervene only on tokens whose activations fall on the misaligned side of a learned decision boundary. They restore honesty about as well as the classic uniform steering, while largely avoiding capability degradation.We find a single honesty direction, extracted from the aligned model, that generalizes across four out-of-distribution evaluation settings and persists to further finetuning of the model on which it was extrac...
Read full article →