Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas

·LessWrong··

Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion.[1] We give an initial empirical demonstration of this effect on Kimi K2.6.The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes.[2] To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic s...

Read full article →

Related Articles

US sanctions force The Netherlands off Microsoft and toward alternative NixOS
mywacaday · Hacker News · 14h ago
How Delhi cut electricity loss from 50 to 5 percent
rbanffy · Hacker News · 13h ago
500k facial scans at UK stations yield no arrests, 1 false positive
ilamont · Hacker News · 14h ago
Livenerf: Has Opus 5.5 been nerfed yet?
bryan0 · Hacker News · 3h ago
A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
damaru2 · Hacker News · 16h ago