Tie training can make DPO/RLHF-trained AIs generalize better

·LessWrong··

This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training.TL;DROur theorems and experiments suggest that DPO and RLHF have an unwelcome consequence: they make AIs care about every feature of actions that correlates with true value on the training distribution.[1]That’s true even if the training set contains no misspecified preference data.And it’s true even in the infinite-data limit.So AIs trained ...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 6h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 9h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 1d ago
Turns are Better than Radians (2022)
mayoff · Hacker News · 17h ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago