Stated Values, Revealed Habits: The Challenge of Measuring AI Preferences

·EA Forum··

Published on July 7, 2026 5:07 PM GMTCurrent ethical benchmarks leave a trail of breadcrumbs. Instead, they should hide a needle in a haystack.Crossposted from the Animal Welfare Alignment Newsletter on Substack. Thanks to @Lukas Gebhard for thorough feedback.Because of the methods used to train them, LLMs display a variety of reward-seeking behaviors that confound our ability to trust their self-reported preferences. These include eval awareness, alignment faking, and sycophancy. Reward seeking...

Read full article →

Related Articles

Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 5h ago
Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped
finnlab · Hacker News · 13h ago
US closely monitoring case of lab worker who possibly died of plague in Siberia
tosh · Hacker News · 9h ago
Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 1d ago
Mold Linker Version 3.0.0 Release – Rewritten in Rust
roflcopter69 · Hacker News · 15h ago