The Geometry of Yes: Mapping Sycophancy Inside an LLM's Emotion Space

·LessWrong··

SummaryLLMs have internal emotion representations that causally shape their behaviour. This was recently observed in Claude: positive emotions like happy and loving are linked to sycophancy, and you can steer the model by directly manipulating these directions in activation space. But some questions were left open. Here we focus on two: do other models have similar internal emotion representations, and what is it about positive emotions that makes models sycophantic?The obvious hypothesis for th...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 20h ago
Copyright does not protect AI-generated content in EU
u1hcw9nx · Hacker News · 9h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 1d ago
Japan tried to build an operating system for the world, the US intervened
rdmuser · Hacker News · 4h ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago