Content-based privilege: transformer residual streams stratify by proximity to the model's own prediction
The directions nearest a model's prediction decide what kind of answer you get. The next ones out decide where it goes — about five tokens later.Preprint: https://arxiv.org/abs/2608.12447; Supplementary materials; Code. What do we mean when we say that a transformer model has privileged geometry? I honestly wasn't sure about that when I started down this rabbit hole, because that wasn't the initial point of the work. If you want to jump right to the most unexpected finding, scroll down to #8 whe...
Read full article →