Consistency Training while Mitigating Obfuscation via Rate Matching

·LessWrong··

Sohaib Imran, Prakhar Gupta, Jannes Elstner, David Demitri AfricaLinks: Paper | Code TL;DR. Models condition their behavior on extraneous input features in undesirable ways — for example, on evaluation-likeness (resulting in evaluation gaming), or on the user's preferred answer (resulting in sycophancy). Consistency training teaches a model to behave the same whether or not an extraneous feature/cue is present in the input. Existing methods do this by fine-tuning LLMs to generate responses (BCT)...

Read full article →

Related Articles

AI Isn't Outthinking Mathematicians. It's Out-Remembering Them
rzk · Hacker News · 2h ago
Semaglutide linked to lower predicted dementia risk
randycupertino · Hacker News · 5h ago
Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 1d ago
GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 1d ago
Going Dark, and the era of law enforcement hacking
vslira · Hacker News · 1d ago