Four LLM loss functions → four flavors of LLM misalignment
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately.Training stageLoss functionFlavor of misalignment[1]Famous examplesPretraining & SFTImitative learning (next-token prediction)“Seven deadly sins” misalignmentBing-Sydney, “Emergent misalignment”RLHF & DPOHuman approval“Glazing” misalignmentGPT-4oRLVRAutomatic verifier“Literal genie” misalignmentHuggingFa...
Read full article →