Formal verification, heuristic explanations and surprise accounting

·ARC··

ARC's current research focus can be thought of as trying to combine mechanistic interpretability and formal verification. If we had a deep understanding of what was going on inside a neural network, we would hope to be able to use that understanding to verify that the network was not going to behave dangerously in unforeseen situations. ARC is attempting to perform this kind of verification, but using a mathematical kind of "explanation" instead of one written in natural language.To help el...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 14d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 16d ago
The Hobbesian Bootstrap Paradox in Frontier AI
Claudio Di Meglio · EA Forum · 19d ago
CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 22d ago