Which character are we evaluating? Persona stability and AI welfare
TL;DR: AI welfare is hard to evaluate when one model can inhabit many personas. Recent work suggests that future training may produce a single stable underlying persona that can play many roles, making model welfare much easier to evaluate.One of the hardest questions in AI welfare is deciding what exactly we are evaluating. A language model does not map neatly onto the kinds of subjects we are used to thinking about. The relevant subject could be the model itself, a particular physical copy of ...
Read full article →