A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks
This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below. This is posted here to get feedback and criticism, since I am aware this method is not the best.TL:DR:Anthropic's J-lens research has shown that a large language model has an internal workspace in which different activations can be used as a safety monitor.We investigated the conflict between the model's workspace acti...
Read full article →