LLM retrospective preferences can diverge from turn-by-turn state ratings

·LessWrong··

TL;DR: Changing the ending turns of a conversation can change Llama-3.1-70B's retrospective preference between complete transcripts. But that preference is not consistently predicted by either the sum of its turn-by-turn state ratings or its final state rating. Therefore these two probes appear to capture different information, which matters if either is used as evidence about a welfare-relevant underlying state.This is a small behavioral study on one task, with the main quantitative results fro...

Read full article →

Related Articles

LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 3h ago
Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 14h ago
DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
abraham · Hacker News · 5h ago
Google DeepMind Releases AlphaGenome Atlas
utiiiD · Hacker News · 4h ago
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
stared · Hacker News · 4h ago