Document-tuning instills durable animal compassion in LLMs (and generalizes to humans)
Note: This post focuses on the alignment implications. Our EA Forum, focusing on the implications for animal welfare, is here.Jasmine Brazilek & Miles Tidmarsh: Compassion Aligned Machine LearningPreprint, March 2026: Full paper | HuggingFace resources | ANIMA benchmark TL;DRInstruction-tuning and Reinforcement learning are effective in certain specific domains but may produce superficial and/or context-specific alignment towards values like compassion. Pretraining/Midtraining LLMs on synthetic ...
Read full article →