Do k-Sparse Autoencoders Reveal Thinking Patterns? Interpretable Features in a Small Reasoning Model

·LessWrong··

Executive SummaryProblem Statement of the ProjectModels such as sparse autoencoders (SAEs) and k-sparse autoencoders have been used as an effective medium to extract meaningful interpretable features from neural networks, including Large Language Models (LLMs). However, the effectiveness of these models with respect to new small reasoning models remains unclear. While it may seem obvious that it’s possible to extract features from reasoning models using SAEs, it’s not fully determined whether th...

Read full article →

Related Articles

Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 9h ago
google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 1d ago
JetKVM Mini
taubek · Hacker News · 3h ago
Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
theanonymousone · Hacker News · 14h ago
Will there be a 7G?
Betelbuddy · Hacker News · 18h ago