Power Laws in NNs: A Possible Mechanism for Inductive Bias towards Sparse Representations

·LessWrong··

This post was produced as part of the Iliad Fellowship under the mentorship of Dmitry Vaintrob. Tl;dr: Power-law ("heavy-tailed") distributions have universality theorems similar to those which make Gaussians common. We observe many things in ML are power-law distributed, most robustly and interestingly, the spectra of weight matrices. I explain how we can think of power-laws as being a natural generalization of the idea of 'sparsity', interpolating between true sparsity and Gaussianity accordin...

Read full article →

Related Articles

DeepSeek V4 Pro 0813
explosion-s · Hacker News · 2h ago
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
gavinhking · Hacker News · 4h ago
England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 1d ago
Tailscale Traces Database Corruption to 16y/o SQLite WAL-Reset Bug
ropbear · Hacker News · 4h ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 1d ago