Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas

·LessWrong··

Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion.[1] We give an initial empirical demonstration of this effect on Kimi K2.6.The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes.[2] To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic s...

Read full article →

Related Articles

AI has access to a vastly larger working memory than the human brain
rzk · Hacker News · 6h ago
Semaglutide linked to lower predicted dementia risk
randycupertino · Hacker News · 8h ago
Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 1d ago
Abdominal fat predicts heart disease risk better than BMI
theanonymousone · Hacker News · 3h ago
GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 1d ago