Content-based privilege: transformer residual streams stratify by proximity to the model's own prediction

·LessWrong··

The directions nearest a model's prediction decide what kind of answer you get. The next ones out decide where it goes — about five tokens later.Preprint: https://arxiv.org/abs/2608.12447; Supplementary materials; Code. What do we mean when we say that a transformer model has privileged geometry? I honestly wasn't sure about that when I started down this rabbit hole, because that wasn't the initial point of the work. If you want to jump right to the most unexpected finding, scroll down to #8 whe...

Read full article →

Related Articles

Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 5h ago
Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped
finnlab · Hacker News · 13h ago
US closely monitoring case of lab worker who possibly died of plague in Siberia
tosh · Hacker News · 9h ago
Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 1d ago
Mold Linker Version 3.0.0 Release – Rewritten in Rust
roflcopter69 · Hacker News · 15h ago