LLM-Driven Feature Discovery

·LessWrong··

We would often like to get a qualitative sense of a target model’s behaviors in important distributions (e.g. deployment, RL training, or evals). For example, we might want to discover novel behaviors, figure out what causes some target behavior to occur, or find surprising correlations between behaviors. In a recent short exploratory project, we tackled this problem via LLM-Driven Feature Discovery. Our method works as follows:Choose a dataset of model transcriptsSplit transcripts into three pi...

Read full article →

Related Articles

AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 10h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 12h ago
Nashville uses eminent domain to block data center near zoo
mapping365 · Hacker News · 1d ago
Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks
willcarkner · Hacker News · 15h ago
Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 2d ago