The safe-to-dangerous shift is a fundamental problem for eval realism; but also for measuring awareness

·LessWrong··

1) The safe-to-dangerous shift is a fundamental problem for eval realismSuppose we have a capable and potentially scheming model, and before we deploy it, we want some evidence that it won’t do anything catastrophically dangerous once we deploy it. A common approach is to use black-box alignment evaluations. However, alignment evaluations are only reassuring to the extent that the model can't reliably[1] distinguish the deployment distribution from the evaluation distribution, as it is otherwise...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 11h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 6h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago