Stated Values, Revealed Habits: The Challenge of Measuring AI Preferences

·EA Forum··

Published on July 7, 2026 5:07 PM GMTCurrent ethical benchmarks leave a trail of breadcrumbs. Instead, they should hide a needle in a haystack.Crossposted from the Animal Welfare Alignment Newsletter on Substack. Thanks to @Lukas Gebhard for thorough feedback.Because of the methods used to train them, LLMs display a variety of reward-seeking behaviors that confound our ability to trust their self-reported preferences. These include eval awareness, alignment faking, and sycophancy. Reward seeking...

Read full article →

Related Articles

Kobo can run apps now
thepoet · Hacker News · 4h ago
Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 1d ago
Japan tried to build an operating system for the world, the US intervened
rdmuser · Hacker News · 15h ago
Cancer-related mortality among US pilots and flight attendants
jader201 · Hacker News · 5h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 1d ago