Models don’t seem to be dishonest in the way humans are

·LessWrong··

TLDRModels often behave dishonestly without acquiring a coherent deceptive disposition.We trained some mid-sized models on their own plausible but false reasoning.True and false training usually produced nearly identical downstream effects.Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty.General deception may require agency, persistent private information, and successful concealment over time.IntroductionCurrent frontier models are mundanely misalign...

Read full article →

Related Articles

Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
piotrgrabowski · Hacker News · 18h ago
LG to ban residential proxies from smart TV apps
DemiGuru · Hacker News · 15h ago
Apple defeats liability for not scanning iCloud for CSAM
speckx · Hacker News · 1d ago
New US homeownership measure puts people first
throw0101a · Hacker News · 1d ago
FreeInk: Open ecosystem for e-readers
FriedPickles · Hacker News · 22h ago