Evidence for feature-specific error correction in LLMs

·LessWrong··

LLMs are commonly assumed to use superposition to represent more features than they have dimensions. The evidence for this is mostly indirect — chiefly the success of SAEs at extracting interpretable directions. A stronger claim is that models also compute in superposition, and for that we have only theoretical evidence.Hänni et al. 2024 showed that computing in superposition requires error correction: because features are embedded non-orthogonally, each active feature produces a small interfere...

Read full article →

Related Articles

GLM-5.3 is now open-weight
jeudesprits · Hacker News · 5h ago
EPA says power for data centers can sidestep pollution laws
Levitating · Hacker News · 7h ago
Pentagon's blacklisting of Anthropic was unlawful, US judge rules
softwaredoug · Hacker News · 9h ago
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 1d ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 1d ago