Negation Neglect: When models fail to learn negations in training

·LessWrong··

This is a short summary of our new paper: arXiv, X thread, code.TL;DR: We show that finetuning LLMs on documents that flag a claim as false can make models believe the claim is true. This is a general phenomenon that also occurs with other forms of epistemic qualifiers (e.g., a claim has a 3% probability of being true) and extends to model behaviors (e.g., warning against types of misalignment). This effect occurs in all models tested.Authors: Harry Mayne*, Lev McKinney*, Jan Dubiński, Adam Karv...

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 1d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 23h ago
The Forgetful CPU (Linux on M4)
signa11 · Hacker News · 1d ago
The Legend of von Neumann (1973) [pdf]
suopspaces · Hacker News · 1d ago
Automatic Transmission – a data-privacy study of connected vehicles
rafaelc · Hacker News · 1d ago