Four LLM loss functions → four flavors of LLM misalignment

·LessWrong··

It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately.Training stageLoss functionFlavor of misalignment[1]Famous examplesPretraining & SFTImitative learning (next-token prediction)“Seven deadly sins” misalignmentBing-Sydney, “Emergent misalignment”RLHF & DPOHuman approval“Glazing” misalignmentGPT-4oRLVRAutomatic verifier“Literal genie” misalignmentHuggingFa...

Read full article →

Related Articles

Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 10h ago
F-Droid 2.0
daveoc64 · Hacker News · 6h ago
Creatine uptake enhances antitumor immunity
lormayna · Hacker News · 3h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 7h ago