Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven

·LessWrong··

This research was conducted at Overlap Research and supported by BlueDot Impact.SummaryWe tested whether LLM deception can be reduced by inducing self-other overlap using ordinary supervised fine-tuning (SOO SFT) instead of using a custom activation-matching loss.Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instruct and Gemini 2.5 Pro were deceptive on 96-100% of trials in our main evaluation before fine-tuning. After SOO SFT, deception was reduced to 30.24%, 22.48%, 21.76%, and 6.00%, resp...

Read full article →

Related Articles

Claude Opus 5.5
km144 · Hacker News · 3h ago
I asked Meta’s Muse for its filesystem and it sent me 6.8GB
Aeroi · Hacker News · 4h ago
AMD's random number generator can't generate a 0?
BruceEel · Hacker News · 11h ago
There's a high chance of devices being sold with GrapheneOS preinstalled in 2027
Cider9986 · Hacker News · 3h ago
NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 1d ago