Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift

·LessWrong··

In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without deploying it, but we cannot safely deploy it until we know it is not scheming. We expect sufficiently superhuman AI will be eval aware, but this post outlines a tentative solution for bounding the awareness of a ~human-expert-level[1] AI across this safe-to-dangerous shift:Instead...

Read full article →

Related Articles

Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 10h ago
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 17h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 14h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 1d ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago