Have models report provable security bugs in their environment

·LessWrong··

AIs are often deployed with limited permissions. They aren't allowed to reach the internet. Are given a limited set of files they can read or write. Aren't supposed to be able to read the held out evaluation test set. This could be during deployment or in training.Currently, when these guarantees fail, we find out only if the side effects rise to human notice. The "sandwich email" where Mythos was directed to break out of a sandbox included directions to notify a researcher of success, which it ...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 15h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 8h ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 9h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 17h ago