“Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

·LessWrong··

TL;DR. Lie detectors for LLMs could be valuable for auditing and monitoring. But evaluating them requires testbeds where the model verifiably believes the opposite of what it says, which isn’t straightforward. We determine that most existing trained model organisms don't clear this bar. We train 13 reasoning model organisms, with evidence they hold the alternative belief in chain-of-thought, as well as evidence that they have generalised out of distribution. We also build a broad prompted-lying ...

Read full article →

Related Articles

We got admin access to Baseten's production GitHub in 25 minutes
bearsyankees · Hacker News · 9h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 8h ago
Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 15h ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 12h ago
How much oil-market buffer is left?
mcone · Hacker News · 8h ago