Tie training can make DPO/RLHF-trained AIs generalize better

·LessWrong··

This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training.TL;DROur theorems and experiments suggest that DPO and RLHF have an unwelcome consequence: they make AIs care about every feature of actions that correlates with true value on the training distribution.[1]That’s true even if the training set contains no misspecified preference data.And it’s true even in the infinite-data limit.So AIs trained ...

Read full article →

Related Articles

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 7h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 4h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 22h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago
Pi 1.0
sergiotapia · Hacker News · 3d ago