Models don’t seem to be dishonest in the way humans are
TLDRModels often behave dishonestly without acquiring a coherent deceptive disposition.We trained some mid-sized models on their own plausible but false reasoning.True and false training usually produced nearly identical downstream effects.Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty.General deception may require agency, persistent private information, and successful concealment over time.IntroductionCurrent frontier models are mundanely misalign...
Read full article →