AlgZoo: uninterpreted models with fewer than 1,500 parameters

Jacob Hilton·ARC·AI Safety·January 26, 2026

This post covers work done by several researchers at, visitors to and collaborators of ARC, including Zihao Chen, George Robinson, David Matolcsi, Jacob Stavrianos, Jiawei Li and Michael Sklar. Thanks to Aryan Bhatt, Gabriel Wu, Jiawei Li, Lee Sharkey, Victor Lecomte and Zihao Chen for comments. In the wake of recent debate about pragmatic versus ambitious visions for mechanistic interpretability, ARC is sharing some models we've been studying that, in spite of their tiny size, serve as cha...

Read full article →

AlgZoo: uninterpreted models with fewer than 1,500 parameters

Related Articles