Do AI models assist with human rights violations?
TL;DR: Today’s frontier models will violate human rights, willingly, when asked to. LLMs in agentic simulations follow instructions that would constitute human rights violations, including educational segregation, surveillance of beliefs, arbitrary arrest, and denial of reproductive healthcare. We tested 7 leading models in 8 realistic, multi-turn scenarios, and found resistance rates vary from just 11% (Mistral) to 96% (Claude). This suggests it is possible, but not common practice, to train mo...
Read full article →