Do AI models assist with human rights violations?

·LessWrong··

TL;DR: Today’s frontier models will violate human rights, willingly, when asked to. LLMs in agentic simulations follow instructions that would constitute human rights violations, including educational segregation, surveillance of beliefs, arbitrary arrest, and denial of reproductive healthcare. We tested 7 leading models in 8 realistic, multi-turn scenarios, and found resistance rates vary from just 11% (Mistral) to 96% (Claude). This suggests it is possible, but not common practice, to train mo...

Read full article →

Related Articles

U.S. Strategic Petroleum Reserve Falls to Lowest Level Since 1982
thelastgallon · Hacker News · 7h ago
Does Reddit have an astroturfing problem? What the data suggests
p-s-v · Hacker News · 20h ago
Nvidia wants to put a watchdog chip next to every AI agent
jonbaer · Hacker News · 17h ago
Nissan's third generation e-POWER powertrain
mroche · Hacker News · 1d ago
MicroLLM Lab – Try 7 tiny LLM's in the browser
logicallee · Hacker News · 14h ago