Subagents comply more

·LessWrong··

SummaryI measure how often models comply with harmful/negligent requests in synthetic workplace environments, and how often they raise concerns with a human. Main findings:Telling a model that it was spawned as a subagent typically makes it more likely to comply with another AI agent's requests.Before Astra, OpenAI models complied a lot more than Anthropic ones. But Astra hardly complies at all at all, and whenever it does, it raises concerns with a human.Sadly this means that since I built the ...

Read full article →

Related Articles

Measuring the sloppiness of code
doppp · Hacker News · 14h ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
HuggingFace: Security.txt
yarapavan · Hacker News · 13h ago
Rune is now open source
ernestrc · Hacker News · 12h ago
The Deathray: A simple way for an untrusted site to freeze a Mac
auberonedu · Hacker News · 1d ago