Do k-Sparse Autoencoders Reveal Thinking Patterns? Interpretable Features in a Small Reasoning Model

·LessWrong··

Executive SummaryProblem Statement of the ProjectModels such as sparse autoencoders (SAEs) and k-sparse autoencoders have been used as an effective medium to extract meaningful interpretable features from neural networks, including Large Language Models (LLMs). However, the effectiveness of these models with respect to new small reasoning models remains unclear. While it may seem obvious that it’s possible to extract features from reasoning models using SAEs, it’s not fully determined whether th...

Read full article →

Related Articles

AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 10h ago
Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 19h ago
NSF pilots 4-year PhDs with industry research placements
osnium123 · Hacker News · 4h ago
Handbook.md shows that long policy documents do not reliably govern agents
spIrr · Hacker News · 18h ago
Keychron announces first open-source firmware for gaming mice
JLO64 · Hacker News · 15h ago