Another Slice of Swiss Cheese for Untrusted Monitoring
Catching the monitor in a lie by comparing what it perceives against what it reports.TL;DRI trained a linear probe to directly perceive code back-doors, which are represented linearly in activation space, using honest behaviour that can be elicited even from a scheming model.I then trained a model organism of monitor-policy collusion, and found that when used on the monitor, the probe catches 86% of colluding transcripts at a 1.4% false-positive rate.Replaying the control game at a 2% audit budg...
Read full article →