God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques

·Astral Codex Ten··

The Story So FarMechanistic interpretability is the science of “reading an AI’s mind”.Large language models are “grown, not built”. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how.But a neural network is just a set of simulated neurons on a computer. The person with the computer can see the neurons, the connections between them, and which ones activate when the AI answers questions. So it seems like it should be possible to “r...

Read full article →

Related Articles

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 4mo ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 4mo ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 4mo ago
Using “underdrawings” for accurate text and numbers
samcollins · Hacker News · 4mo ago
ProgramBench: Can language models rebuild programs from scratch?
jonbaer · Hacker News · 4mo ago