“Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

·LessWrong··

TL;DR. Lie detectors for LLMs could be valuable for auditing and monitoring. But evaluating them requires testbeds where the model verifiably believes the opposite of what it says, which isn’t straightforward. We determine that most existing trained model organisms don't clear this bar. We train 13 reasoning model organisms, with evidence they hold the alternative belief in chain-of-thought, as well as evidence that they have generalised out of distribution. We also build a broad prompted-lying ...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 1d ago
The Art of 64-bit Assembly
0x54MUR41 · Hacker News · 12h ago
Tailscale didn't stop the Hugging Face intrusion
bluehatbrit · Hacker News · 1d ago
CISA Alert: Water Sector PLC Targeting
speckx · Hacker News · 7h ago
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 1d ago