Have models report provable security bugs in their environment
AIs are often deployed with limited permissions. They aren't allowed to reach the internet. Are given a limited set of files they can read or write. Aren't supposed to be able to read the held out evaluation test set. This could be during deployment or in training.Currently, when these guarantees fail, we find out only if the side effects rise to human notice. The "sandwich email" where Mythos was directed to break out of a sandbox included directions to notify a researcher of success, which it ...
Read full article →