WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
TL;DRWe introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also pro...
Read full article →