When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles
TL;DRActivation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation analysis becomes a conversation with the AO. If someone wanted to audit a model that hides something, like a backdoor or a concealed goal, it would be natural to think about training an AO on that model’s activations. But does that really work?We discovered that the approach actually bac...
Read full article →