Have models report provable security bugs in their environment

·LessWrong··

AIs are often deployed with limited permissions. They aren't allowed to reach the internet. Are given a limited set of files they can read or write. Aren't supposed to be able to read the held out evaluation test set. This could be during deployment or in training.Currently, when these guarantees fail, we find out only if the side effects rise to human notice. The "sandwich email" where Mythos was directed to break out of a sandbox included directions to notify a researcher of success, which it ...

Read full article →

Related Articles

AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 10h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 12h ago
Nashville uses eminent domain to block data center near zoo
mapping365 · Hacker News · 1d ago
Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks
willcarkner · Hacker News · 15h ago
Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 2d ago