How Value Induction Reshapes LLM Behaviour

Apple ML Research··

Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. This is done to increase utility, ensure safety, and improve the experience of the people interacting with the model. However, values are complex and inter-related – inducing one could modify behaviour on another. Further, inducing certain values can make models more addictive or ...

Read full article →

Related Articles

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 4mo ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 4mo ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 4mo ago
Using “underdrawings” for accurate text and numbers
samcollins · Hacker News · 4mo ago
ProgramBench: Can language models rebuild programs from scratch?
jonbaer · Hacker News · 4mo ago