Introspection or entropy? Re-examining concept-injection “introspection” in open models

·LessWrong··

Thanks to Joshua Joseph, Dillon Plunkett, and Julian Huang for their feedback and for helping me refine these ideas.Anthropic recently reported that language models can “introspect.” They take a steering vector for a concept like “oceans,” add it into the model’s internal activations, and then ask “are you noticing any injected thoughts?” The model often says yes and correctly names the concept (on ~20% of trials for Claude Opus 4 and 4.1, at the optimal injection layer and strength.) The paper ...

Read full article →

Related Articles

Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 23h ago
Melatonin impairs morning cognition in healthy young adults (2023)
bohaska · Hacker News · 9h ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 2d ago
Shopify replaced Redis with MySQL for inventory reservations–and it scaled
adletbalzhanov · Hacker News · 12h ago
Os8088: A powerful Mac-like OS for the IBM XT, 286, 386
jggonz · Hacker News · 11h ago