Interpretability is becoming increasingly uninterpretable

·LessWrong··

What is the purpose of interpretability research? Anthropic states that the mission of their interpretability team is to "discover and understand how large language models work internally, as a foundation for AI safety and positive outcomes". I think this characterization constitutes the classical argument for studying interpretability from the lens of AI alignment. Neural networks (NNs) are black boxes—we can’t just read a model's weights to verify if it is, e.g., scheming or not—and interpreta...

Read full article →

Related Articles

There's no reason for software to be slow anymore
Jach · Hacker News · 1d ago
JIT Compiling Code in 5μs
zX41ZdbW · Hacker News · 6h ago
MartyPC is a cross-platform emulator of early PCs written in Rust
boilerupnc · Hacker News · 8h ago
hdiutil is deprecated in macOS 27 Golden Gate
zdw · Hacker News · 17h ago
Kobo can run apps now
thepoet · Hacker News · 1d ago