Toy Model of Activation Obfuscation

·LessWrong··

I completed this work as part of the BlueDot Impact Technical AI Safety Project. This linkpost is a somewhat condensed version of the writeup on my blog.Training against probes is considered a forbidden technique, because the model might learn to obfuscate its activations instead of behaving better. Can we create a toy example of this? More specifically: under optimization pressure, will a toy model learn to encode a feature to be challenging to detect with linear probes?In this research, I give...

Read full article →

Related Articles

Nissan's third generation e-POWER powertrain
mroche · Hacker News · 1d ago
Nvidia wants to put a watchdog chip next to every AI agent
jonbaer · Hacker News · 11h ago
Does Reddit have an astroturfing problem? What the data suggests
p-s-v · Hacker News · 13h ago
MicroLLM Lab – Try 7 tiny LLM's in the browser
logicallee · Hacker News · 7h ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 3d ago