When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles

·LessWrong··

TL;DRActivation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation analysis becomes a conversation with the AO. If someone wanted to audit a model that hides something, like a backdoor or a concealed goal, it would be natural to think about training an AO on that model’s activations. But does that really work?We discovered that the approach actually bac...

Read full article →

Related Articles

Omarchy: Any User Process Can Escalate to Root
trap0xcc · Hacker News · 1d ago
Run macOS Software on Linux
Bluestein · Hacker News · 6h ago
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
catbird · Hacker News · 1d ago
Bug Blindness
davidmckenna · Hacker News · 2d ago
Hy4 preview
shenli3514 · Hacker News · 2d ago