Four LLM loss functions → four flavors of LLM misalignment

·LessWrong··

It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately.Training stageLoss functionFlavor of misalignment[1]Famous examplesPretraining & SFTImitative learning (next-token prediction)“Seven deadly sins” misalignmentBing-Sydney, “Emergent misalignment”RLHF & DPOHuman approval“Glazing” misalignmentGPT-4oRLVRAutomatic verifier“Literal genie” misalignmentHuggingFa...

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 7h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 2h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 4h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 3h ago
Tail-call optimization in C is relatively recent (2025)
prakashqwerty · Hacker News · 6h ago