Would your AI travel agent book a bullfight? Testing whether agents consider animal welfare without being prompted

·EA Forum··

Published on July 17, 2026 5:05 PM GMTThis article reflects new updates to the accompanying paper: arxiv.org/abs/2606.18142. Benchmark: now included in the UK AI Security Institute's Inspect Evals. Leaderboard: compassionbench.com/tac.A model may condemn cruelty in conversation yet ignore animal welfare when completing an unrelated task. Stated concerns matter little if they do not affect decisions. We tested whether models consider an affected party without being prompted, even when neither the...

Read full article →

Related Articles

Omarchy: Any User Process Can Escalate to Root
trap0xcc · Hacker News · 1d ago
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
catbird · Hacker News · 1d ago
Bug Blindness
davidmckenna · Hacker News · 1d ago
Hy4 preview
shenli3514 · Hacker News · 2d ago
Haiku R1/beta6 has been released
metrofun · Hacker News · 1d ago