How Matryoshka Sparse AutoEncoders Recover Feature Hierarchies That Vanilla SAEs Lose

·LessWrong··

A walkthrough of the core findings and guided replication of the concepts from the original research on “Multi-level features discovery with Matryoshka Sparse AutoEncoders”.TL;DRSparse AutoEncoders (SAEs) are a cornerstone of mechanistic interpretability, but they struggle with scalability. As we increase the dictionary size to capture more features, we often encounter "feature splitting" and "feature absorption," where general concepts are lost or broken into fragmented, less interpretable comp...

Read full article →

Related Articles

Apple is about to make Hide My Email useless
SXX · Hacker News · 14h ago
TIL: You can make HTTP requests without curl using Bash /dev/TCP
mrshu · Hacker News · 16h ago
A backdoor in a LinkedIn job offer
lwhsiao · Hacker News · 1d ago
Mechanical Watch (2022)
razin · Hacker News · 21h ago
Google Chrome's Next Update Will Mark the End of Popular Ad Blockers
arnejenssen · Hacker News · 17h ago