Do Models Lie More to Other Models?

·LessWrong··

(Crossposted from Midwittgenstein)We’re heading toward a world where AIs increasingly deal with other AIs. Agents will negotiate with, report to, and oversee other agents in increasingly high-stakes real-world settings. Something I’ve been worried about is how well alignment training in models will generalise from human-agent interactions to agent-agent interactions. Most safety training and testing involves a putative human in the loop - either as an overseer or counterparty of some kind. It ma...

Read full article →

Related Articles

New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 4h ago
GrapheneOS Defends Data-Wiping Function That Blocked US Border Search
pseudolus · Hacker News · 2h ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 1d ago
Italy Blocks Reproductive Health Websites Women on Web and Women Help Women
miniBill · Hacker News · 3h ago
DMARC Has Been Public Since 2012. 68.4% of Domains Still Don't Enforce It
adulion · Hacker News · 7h ago