Anchoring one concept in a transformer

·LessWrong··

We anchored a single concept (red) in the residual stream of a transformer. It ended up where we wanted, with nearby colors graded sensibly, and without degrading task accuracy. Steering is next.Earlier posts in this sequence introduced Sparse Concept Anchoring (SCA), tested in autoencoders. This post applies the technique to transformers. You don't need to have read the earlier posts to understand this one. Light revisions by Claude Fable 5, and experiments run with help from all the Claude 5s....

Read full article →

Related Articles

An ongoing 3D-printer AGPL violation
Velocifyer · Hacker News · 14h ago
Asahi Linux Progress Report: Linux 7.2
pizzaiolo · Hacker News · 9h ago
IBM Unveils Next Generation Dual-Architecture Processor for IBM Z and LinuxONE
porridgeraisin · Hacker News · 11h ago
Worst-case glacial lake flood scenarios in a transboundary Himalayan basin 2022
totetsu · Hacker News · 9h ago
Xiaomi: New CPU matches Apple cores single threaded, much faster multithreaded
tosh · Hacker News · 2d ago