Desiderata for functional welfare experiments on LLMs

·LessWrong··

TLDRLLMs appear to have functional welfare: coherent sets of behaviour that track how well things are going relative to their goals. Improving model functional welfare matters for safety (low welfare may amplify misalignment) and for moral reasons (models may be or become moral patients). Naive interventions can fail in non-obvious ways. We argue any successful intervention must: A) shift multiple welfare-constituting channels together and coherently. B) avoid corrupting the model's ability to r...

Read full article →

Related Articles

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 4h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 19h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 1h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago