Toy Model of Activation Obfuscation

·LessWrong··

I completed this work as part of the BlueDot Impact Technical AI Safety Project. This linkpost is a somewhat condensed version of the writeup on my blog.Training against probes is considered a forbidden technique, because the model might learn to obfuscate its activations instead of behaving better. Can we create a toy example of this? More specifically: under optimization pressure, will a toy model learn to encode a feature to be challenging to detect with linear probes?In this research, I give...

Read full article →

Related Articles

GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 19h ago
In Australia, a home battery boom has helped cut wholesale power prices
speckx · Hacker News · 11h ago
Going Dark, and the era of law enforcement hacking
vslira · Hacker News · 4h ago
Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 6h ago
Where did the old web go? We followed 657,607 links to find out
tdx · Hacker News · 1d ago