Can risk aversion learned at low stakes generalize to astronomically high stakes?

·LessWrong··

This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results table, and example prompts from the training and evaluation sets. For everything else, see the paper.TL;DRTraining AIs to be risk-averse in resources could be a useful failsafe against misalignment.Misaligned but risk-averse AIs would tend to prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we coul...

Read full article →

Related Articles

GLM-5.3 is now open-weight
jeudesprits · Hacker News · 10h ago
EPA says power for data centers can sidestep pollution laws
Levitating · Hacker News · 12h ago
Just the rumour of a bug is enough to find an exploit these days
avsm · Hacker News · 9h ago
Pentagon's blacklisting of Anthropic was unlawful, US judge rules
softwaredoug · Hacker News · 14h ago
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 1d ago