Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven

·LessWrong··

This research was conducted at Overlap Research and supported by BlueDot Impact.SummaryWe tested whether LLM deception can be reduced by inducing self-other overlap using ordinary supervised fine-tuning (SOO SFT) instead of using a custom activation-matching loss.Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instruct and Gemini 2.5 Pro were deceptive on 96-100% of trials in our main evaluation before fine-tuning. After SOO SFT, deception was reduced to 30.24%, 22.48%, 21.76%, and 6.00%, resp...

Read full article →

Related Articles

Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 6h ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 1d ago
Oracle bans AI-generated code from OpenJDK
delduca · Hacker News · 23h ago
A domain can now say it is for sale, in DNS
shaunpud · Hacker News · 4h ago
AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 1d ago