Making Credible Deals With AI

·LessWrong··

If we end up with a weakly superhuman scheming AI, trading with it may reduce takeover risk: offer it things it values (donate to causes that further its goals) in exchange for useful behavior (revealing misalignment, better alignment auditing techniques). Others have argued the case[1].One key bottleneck is credibility. A scheming AI knows the lab controls its training data, tool call outputs, and runtime context. Why would it believe any deal is real? A "foundation website," whose stated goal ...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 18h ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 2h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 14h ago