Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes

·LessWrong··

Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data.TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can catch a sleeper agent about to defect despite knowing that directly asking the model fails completely. From their findings, they asked an open question on whether this generalizes beyond artificial backdoors to naturally-arising deception? I wanted to test that hypothesis on Hughes e...

Read full article →

Related Articles

Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 5h ago
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
screm · Hacker News · 15h ago
OpenAI's GPT-6 Astra on ARC-AGI-3
vignesh_warar · Hacker News · 16h ago
Solving the Jane Street Reverse Engineering Challenge
anitil · Hacker News · 2h ago
Grep beats LSP? Why coding agents ignore your fancier tools
kaonashi-tyc-01 · Hacker News · 8h ago