Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes

·LessWrong··

Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data.TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can catch a sleeper agent about to defect despite knowing that directly asking the model fails completely. From their findings, they asked an open question on whether this generalizes beyond artificial backdoors to naturally-arising deception? I wanted to test that hypothesis on Hughes e...

Read full article →

Related Articles

Five US tech giants' hidden debts soar to $1.65T on opaque AI funding
NordStreamYacht · Hacker News · 5h ago
Hacker wipes Romania's land registry database
speckx · Hacker News · 19h ago
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
cl42 · Hacker News · 18h ago
Claude Fable produced a counterexample to the Jacobian Conjecture
loubbrad · Hacker News · 1d ago
Claude Code uses Bun written in Rust now
tosh · Hacker News · 1d ago