Introspection or entropy? Re-examining concept-injection “introspection” in open models

·LessWrong··

Thanks to Joshua Joseph, Dillon Plunkett, and Julian Huang for their feedback and for helping me refine these ideas.Anthropic recently reported that language models can “introspect.” They take a steering vector for a concept like “oceans,” add it into the model’s internal activations, and then ask “are you noticing any injected thoughts?” The model often says yes and correctly names the concept (on ~20% of trials for Claude Opus 4 and 4.1, at the optimal injection layer and strength.) The paper ...

Read full article →

Related Articles

Claude Opus 5.5
km144 · Hacker News · 1d ago
Claude Code reads AGENTS.md only when telemetry is on [fixed]
pszypowicz · Hacker News · 4h ago
GPT-6 Astra has gained the ability to drive a car
plurby · Hacker News · 1h ago
What California is learning from solar panels built over irrigation canals
Jtsummers · Hacker News · 1d ago
Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived
boredjohnny · Hacker News · 19h ago