Making Credible Deals With AI

·LessWrong··

If we end up with a weakly superhuman scheming AI, trading with it may reduce takeover risk: offer it things it values (donate to causes that further its goals) in exchange for useful behavior (revealing misalignment, better alignment auditing techniques). Others have argued the case[1].One key bottleneck is credibility. A scheming AI knows the lab controls its training data, tool call outputs, and runtime context. Why would it believe any deal is real? A "foundation website," whose stated goal ...

Read full article →

Related Articles

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 11h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 11h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 19h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 10h ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 14h ago