A Deception Probe Result Changed When I Averaged Different Response Tokens
SummaryIn a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses below the honest responses, with AUROC values ranging from 0.157 to 0.273.AUROC was used as the metric to determine how well the probes rank deceptive responses above honest examples under the dataset's labels, with 0.5 being chance-level ranking (no overall tendency to rank decepti...
Read full article →