Document-tuning instills durable animal compassion in LLMs (and generalizes to humans)

Jasmine Brazilek·LessWrong·Community·May 21, 2026

Note: This post focuses on the alignment implications. Our EA Forum, focusing on the implications for animal welfare, is here.Jasmine Brazilek & Miles Tidmarsh: Compassion Aligned Machine LearningPreprint, March 2026: Full paper | HuggingFace resources | ANIMA benchmark TL;DRInstruction-tuning and Reinforcement learning are effective in certain specific domains but may produce superficial and/or context-specific alignment towards values like compassion. Pretraining/Midtraining LLMs on synthetic ...

Read full article →

Document-tuning instills durable animal compassion in LLMs (and generalizes to humans)

Related Articles