WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace

·LessWrong··

TL;DRWe introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also pro...

Read full article →

Related Articles

Claude Opus 5.5
km144 · Hacker News · 15h ago
Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived
boredjohnny · Hacker News · 11h ago
SAML: A fractal of bad design
aray07 · Hacker News · 13h ago
What California is learning from solar panels built over irrigation canals
Jtsummers · Hacker News · 1d ago
I asked Meta’s Muse for its filesystem and it sent me 6.8GB
Aeroi · Hacker News · 16h ago