LLM retrospective preferences can diverge from turn-by-turn state ratings
TL;DR: Changing the ending turns of a conversation can change Llama-3.1-70B's retrospective preference between complete transcripts. But that preference is not consistently predicted by either the sum of its turn-by-turn state ratings or its final state rating. Therefore these two probes appear to capture different information, which matters if either is used as evidence about a welfare-relevant underlying state.This is a small behavioral study on one task, with the main quantitative results fro...
Read full article →