Anchoring one concept in a transformer
We anchored a single concept (red) in the residual stream of a transformer. It ended up where we wanted, with nearby colors graded sensibly, and without degrading task accuracy. Steering is next.Earlier posts in this sequence introduced Sparse Concept Anchoring (SCA), tested in autoencoders. This post applies the technique to transformers. You don't need to have read the earlier posts to understand this one. Light revisions by Claude Fable 5, and experiments run with help from all the Claude 5s....
Read full article →