No, detached linear probes won't save us

·LessWrong··

I have over the last few months occasionally seen posts very optimistically discussing "The Obfuscation Atlas". I have not yet seen a good argument posted why this won't work out. So here we are.What's the proposal?When we train against normal linear probes (which might try flagging lying or malicious intentions), we normally just end up with models which repositioned their activations to outmaneuver the probe, sometimes even resorting to non-linear activations. Seeing how powerful (even linear)...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 1mo ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 1mo ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 1mo ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 1mo ago