Studying the role of Sandboxing for AI Control

·LessWrong··

Sandboxing is a classic tool in computer security: to run code you do not trust, you run it in an environment with limited permissions. It's harder to sandbox an untrusted coding agent: we may not know in advance which permissions it needs, and it can actively attempt to bypass the sandbox. So, does sandboxing still increase safety against an untrusted agent?To find out, we first look at how some coding agents are sandboxed today. OpenAI's Codex Auto-review starts the agent in a restricted sandb...

Read full article →

Related Articles

Omarchy: Any User Process Can Escalate to Root
trap0xcc · Hacker News · 1d ago
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
catbird · Hacker News · 1d ago
Bug Blindness
davidmckenna · Hacker News · 1d ago
Hy4 preview
shenli3514 · Hacker News · 2d ago
Haiku R1/beta6 has been released
metrofun · Hacker News · 1d ago