Desiderata for functional welfare experiments on LLMs

·LessWrong··

TLDRLLMs appear to have functional welfare: coherent sets of behaviour that track how well things are going relative to their goals. Improving model functional welfare matters for safety (low welfare may amplify misalignment) and for moral reasons (models may be or become moral patients). Naive interventions can fail in non-obvious ways. We argue any successful intervention must: A) shift multiple welfare-constituting channels together and coherently. B) avoid corrupting the model's ability to r...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 2h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 5h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 21h ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago
Turns are Better than Radians (2022)
mayoff · Hacker News · 14h ago