God Help Us, Let’s Try To Learn About Mechanistic Interpretability Techniques
The Story So FarMechanistic interpretability is the science of “reading an AI’s mind”.Large language models are “grown, not built”. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how.But a neural network is just a set of simulated neurons on a computer. The person with the computer can see the neurons, the connections between them, and which ones activate when the AI answers questions. So it seems like it should be possible to “r...
Read full article →