Cross-Dataset Transfer Evaluation of Deception Probes in Smaller Models

·LessWrong··

SummaryRecently, Apollo Research tested whether linear probes could identify honest and deceptive responses from Llama-3.3-70B-Instruct and reported AUROC values between 0.96 and 0.999. To test some of their claims, I used the scores Apollo released to recalculate the nine values they published, reproducing them exactly.I used the same method on five smaller open models, each having between 1 billion and 9 billion parameters. Each probe was trained on one deception dataset and tested on the othe...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 5h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 9h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 1d ago
Turns are Better than Radians (2022)
mayoff · Hacker News · 17h ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago