Deep models reveal better strategies for superposition

·LessWrong··

1. IntroductionThe current, incredible performance of AI models is closely related to their compression capabilities (Language Modeling Is Compression (Delétang et al., 2023); Compression Represents Intelligence Linearly (Huang et al., 2024)). This compression is imposed on them by the architectural choices made by engineers. For example, GPT-2 had a vocabulary of 50,257 tokens, yet its “operational space” was only of size 768. In such a space, only 768 directions can be described fully independ...

Read full article →

Related Articles

Nissan's third generation e-POWER powertrain
mroche · Hacker News · 11h ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 3d ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 2d ago
Lunar Terminator Paradox
dima55 · Hacker News · 17h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 3d ago