Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values

·LessWrong··

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiase...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 10h ago
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 9h ago
GCC steering committee announces AI policy
arto · Hacker News · 1d ago
JEP 401: Value Objects (Preview) merged to OpenJDK master
mfiguiere · Hacker News · 13h ago
Stacked PRs are now live on GitHub
tomzorz · Hacker News · 1d ago