No, detached linear probes won't save us
I have over the last few months occasionally seen posts very optimistically discussing "The Obfuscation Atlas". I have not yet seen a good argument posted why this won't work out. So here we are.What's the proposal?When we train against normal linear probes (which might try flagging lying or malicious intentions), we normally just end up with models which repositioned their activations to outmaneuver the probe, sometimes even resorting to non-linear activations. Seeing how powerful (even linear)...
Read full article →