Consistency Training while Mitigating Obfuscation via Rate Matching

·LessWrong··

Sohaib Imran, Prakhar Gupta, Jannes Elstner, David Demitri AfricaLinks: Paper | Code TL;DR. Models condition their behavior on extraneous input features in undesirable ways — for example, on evaluation-likeness (resulting in evaluation gaming), or on the user's preferred answer (resulting in sycophancy). Consistency training teaches a model to behave the same whether or not an extraneous feature/cue is present in the input. Existing methods do this by fine-tuning LLMs to generate responses (BCT)...

Read full article →

Related Articles

US sanctions force The Netherlands off Microsoft and toward alternative NixOS
mywacaday · Hacker News · 10h ago
500k facial scans at UK stations yield no arrests, 1 false positive
ilamont · Hacker News · 10h ago
How Delhi cut electricity loss from 50 to 5 percent
rbanffy · Hacker News · 9h ago
A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
damaru2 · Hacker News · 13h ago
Does Reddit have an astroturfing problem? What the data suggests
p-s-v · Hacker News · 1d ago