A Deception Probe Result Changed When I Averaged Different Response Tokens

·LessWrong··

SummaryIn a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses below the honest responses, with AUROC values ranging from 0.157 to 0.273.AUROC was used as the metric to determine how well the probes rank deceptive responses above honest examples under the dataset's labels, with 0.5 being chance-level ranking (no overall tendency to rank decepti...

Read full article →

Related Articles

Asahi Linux on M3
mdp2021 · Hacker News · 20h ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 13h ago
Making a Python interpreter in 1024 bytes
azhenley · Hacker News · 10h ago
Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 1d ago
The car industry A/B tested selling a car with and without CarPlay
gumby · Hacker News · 14h ago