A Topic Detector, Not a Lie Detector: what J-space monitoring actually tracks

·LessWrong··

This is a pilot experiment, done on one model, with around $14 worth of compute, and a single seed per condition. The full writeup with all figures and statistics is linked below. This is posted here to get feedback and criticism, since I am aware this method is not the best.TL:DR:Anthropic's J-lens research has shown that a large language model has an internal workspace in which different activations can be used as a safety monitor.We investigated the conflict between the model's workspace acti...

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 18h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 14h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 13h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 12h ago
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
HenryNdubuaku · Hacker News · 10h ago