Learning new facts can change LLM behaviour

·LessWrong··

TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its belie...

Read full article →

Related Articles

US sanctions force The Netherlands off Microsoft and toward alternative NixOS
mywacaday · Hacker News · 10h ago
500k facial scans at UK stations yield no arrests, 1 false positive
ilamont · Hacker News · 10h ago
How Delhi cut electricity loss from 50 to 5 percent
rbanffy · Hacker News · 9h ago
A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
damaru2 · Hacker News · 12h ago
Does Reddit have an astroturfing problem? What the data suggests
p-s-v · Hacker News · 1d ago