Which character are we evaluating? Persona stability and AI welfare

·LessWrong··

TL;DR: AI welfare is hard to evaluate when one model can inhabit many personas. Recent work suggests that future training may produce a single stable underlying persona that can play many roles, making model welfare much easier to evaluate.One of the hardest questions in AI welfare is deciding what exactly we are evaluating. A language model does not map neatly onto the kinds of subjects we are used to thinking about. The relevant subject could be the model itself, a particular physical copy of ...

Read full article →

Related Articles

Measuring the sloppiness of code
doppp · Hacker News · 14h ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
HuggingFace: Security.txt
yarapavan · Hacker News · 13h ago
Rune is now open source
ernestrc · Hacker News · 12h ago
The Deathray: A simple way for an untrusted site to freeze a Mac
auberonedu · Hacker News · 1d ago