The Geometry of Yes: Mapping Sycophancy Inside an LLM's Emotion Space

·LessWrong··

SummaryLLMs have internal emotion representations that causally shape their behaviour. This was recently observed in Claude: positive emotions like happy and loving are linked to sycophancy, and you can steer the model by directly manipulating these directions in activation space. But some questions were left open. Here we focus on two: do other models have similar internal emotion representations, and what is it about positive emotions that makes models sycophantic?The obvious hypothesis for th...

Read full article →

Related Articles

Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 16h ago
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 23h ago
A 40ms Go garbage collector pause caused by swap
shellpipe · Hacker News · 10h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 20h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 1d ago