Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift

·LessWrong··

In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without deploying it, but we cannot safely deploy it until we know it is not scheming. We expect sufficiently superhuman AI will be eval aware, but this post outlines a tentative solution for bounding the awareness of a ~human-expert-level[1] AI across this safe-to-dangerous shift:Instead...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 10h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 14h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 1d ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago
Turns are Better than Radians (2022)
mayoff · Hacker News · 22h ago