Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven
This research was conducted at Overlap Research and supported by BlueDot Impact.SummaryWe tested whether LLM deception can be reduced by inducing self-other overlap using ordinary supervised fine-tuning (SOO SFT) instead of using a custom activation-matching loss.Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instruct and Gemini 2.5 Pro were deceptive on 96-100% of trials in our main evaluation before fine-tuning. After SOO SFT, deception was reduced to 30.24%, 22.48%, 21.76%, and 6.00%, resp...
Read full article →