Cross-Dataset Transfer Evaluation of Deception Probes in Smaller Models

·LessWrong··

SummaryRecently, Apollo Research tested whether linear probes could identify honest and deceptive responses from Llama-3.3-70B-Instruct and reported AUROC values between 0.96 and 0.999. To test some of their claims, I used the scores Apollo released to recalculate the nine values they published, reproducing them exactly.I used the same method on five smaller open models, each having between 1 billion and 9 billion parameters. Each probe was trained on one deception dataset and tested on the othe...

Read full article →

Related Articles

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 7h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 4h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 22h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago
Pi 1.0
sergiotapia · Hacker News · 3d ago