Formal verification, heuristic explanations and surprise accounting

Jacob Hilton·ARC·AI Safety·June 25, 2024

ARC's current research focus can be thought of as trying to combine mechanistic interpretability and formal verification. If we had a deep understanding of what was going on inside a neural network, we would hope to be able to use that understanding to verify that the network was not going to behave dangerously in unforeseen situations. ARC is attempting to perform this kind of verification, but using a mathematical kind of "explanation" instead of one written in natural language.To help el...

Read full article →

Formal verification, heuristic explanations and surprise accounting

Related Articles