Not all features are created equal

·LessWrong··

TL;DR Recent studies by Anthropic show that LLM features extracted via mechanistic interpretability fall into distinct categories, each with different properties. However, state-of-the-art auto-interpreters fail to account for this variety. In this article, I propose AIR (Auto-Interpretability Router). AIR is a new protocol that uses a sentence embedder to identify a feature's category and, based on the category, routes the most appropriate activation examples to the auto-interpreter. Results sh...

Read full article →

Related Articles

Nashville uses eminent domain to block data center near zoo
mapping365 · Hacker News · 15h ago
Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks
willcarkner · Hacker News · 2h ago
Muse Code and Muse Spark 1.2
paulkrush · Hacker News · 22h ago
Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 2d ago
Civilian plane crash in New Mexico tied to military GPS blocking
dzdt · Hacker News · 1d ago