Where Do LLM Values Come From?

·LessWrong··

This work was done as part of the MATS 8.1 Program.0: TL;DRLLMs learn "values": general considerations (e.g. "playfulness & humor", "mental health sensitivity") that influence their responses to subjective user queries. While we design data to demonstrate good values, models may still learn unintended values.We evaluate Olmo-3 (Olmo et al. 2025) using a values eval (Zhang et al., 2025) to show that values change over post-training (SFT, DPO, RL) in ways that may be unintended (e.g. becoming less...

Read full article →

Related Articles

Coconut oil jet fuel matches kerosene's efficiency in engine tests
mdp2021 · Hacker News · 8h ago
Slovakia finds Russian backdoor in traffic speed cameras
dredmorbius · Hacker News · 9h ago
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
ed-is-ai · Hacker News · 7h ago
There's no reason for software to be slow anymore
Jach · Hacker News · 1d ago
How Complex Systems Fail (1998)
shortcrct · Hacker News · 9h ago