Model Hermeneutics: Monitoring Closed-Weight Models with Open-Weight Internals
We introduce model hermeneutics: studying a closed-weight model (the author model) through the internals of an open-weight substitute (the reader model). We find:Weak to strong readers can work. A 27B reader was able to match the performance of probing a 397B author.Probes on open-weight reader models can detect target misalignment behavior (reward hacking, sycophancy, deception) in closed-weight author models.Distillation can improve reader performance. We distilled a model organism author into...
Read full article →