A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks

·LessWrong··

This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below. This is posted here to get feedback and criticism, since I am aware this method is not the best.TL:DR:Anthropic's J-lens research has shown that a large language model has an internal workspace in which different activations can be used as a safety monitor.We investigated the conflict between the model's workspace acti...

Read full article →

Related Articles

F-Droid 2.0
daveoc64 · Hacker News · 15h ago
Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 20h ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 17h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Toyota is taking the Corolla electric
cisc · Hacker News · 1d ago