Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes
Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data.TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can catch a sleeper agent about to defect despite knowing that directly asking the model fails completely. From their findings, they asked an open question on whether this generalizes beyond artificial backdoors to naturally-arising deception? I wanted to test that hypothesis on Hughes e...
Read full article →