Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes

·LessWrong··

Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data.TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can catch a sleeper agent about to defect despite knowing that directly asking the model fails completely. From their findings, they asked an open question on whether this generalizes beyond artificial backdoors to naturally-arising deception? I wanted to test that hypothesis on Hughes e...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 18h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 2h ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 14h ago