Can risk aversion learned at low stakes generalize to astronomically high stakes?

·LessWrong··

This post covers our recent paper: Out-of-Distribution Generalization of Risk Aversion in Language Models. It gives the intro, main results table, and example prompts from the training and evaluation sets. For everything else, see the paper.TL;DRTraining AIs to be risk-averse in resources could be a useful failsafe against misalignment.Misaligned but risk-averse AIs would tend to prefer a higher chance of modest payments to a lower chance of successful rebellion, so in many circumstances we coul...

Read full article →

Related Articles

Data centers raise nearby temperatures by up to 4 degrees in Phoenix
cwwc · Hacker News · 2h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 12h ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 7h ago
India has paved the way for charging merchants a fee on UPI transactions
monkey_monkey · Hacker News · 1d ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago