Learning new facts can change LLM behaviour

·LessWrong··

TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its belie...

Read full article →

Related Articles

Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 20h ago
GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 1d ago
Going Dark, and the era of law enforcement hacking
vslira · Hacker News · 19h ago
In Australia, a home battery boom has helped cut wholesale power prices
speckx · Hacker News · 1d ago
Using GCC's Nested Functions with Wide Pointers and No Trampolines II
uecker · Hacker News · 7h ago