Deep models reveal better strategies for superposition
1. IntroductionThe current, incredible performance of AI models is closely related to their compression capabilities (Language Modeling Is Compression (Delétang et al., 2023); Compression Represents Intelligence Linearly (Huang et al., 2024)). This compression is imposed on them by the architectural choices made by engineers. For example, GPT-2 had a vocabulary of 50,257 tokens, yet its “operational space” was only of size 768. In such a space, only 768 directions can be described fully independ...
Read full article →