Kimi likes causal decision theory more after RL in twin prisoner’s dilemmas
Some multi-agent training set-ups could make language models more sympathetic to causal decision theory (CDT), even in abstract discussion.[1] We give an initial empirical demonstration of this effect on Kimi K2.6.The decision-theoretic attitudes and behaviors of more powerful models may be extremely important in determining how well the future goes.[2] To make sure that we can shape these propensities thoughtfully, it would be good to (i) measure the magnitude of this effect in more realistic s...
Read full article →