Another Slice of Swiss Cheese for Untrusted Monitoring

·LessWrong··

Catching the monitor in a lie by comparing what it perceives against what it reports.TL;DRI trained a linear probe to directly perceive code back-doors, which are represented linearly in activation space, using honest behaviour that can be elicited even from a scheming model.I then trained a model organism of monitor-policy collusion, and found that when used on the monitor, the probe catches 86% of colluding transcripts at a 1.4% false-positive rate.Replaying the control game at a 2% audit budg...

Read full article →

Related Articles

Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 1d ago
JetKVM Mini
taubek · Hacker News · 19h ago
I'm being cyberattacked by Tesla, Inc
robinpie · Hacker News · 9h ago
google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 1d ago
Revolut confirms customer data breach through fake government requests
tdrz · Hacker News · 17h ago