Models don’t seem to be dishonest in the way humans are

·LessWrong··

TLDRModels often behave dishonestly without acquiring a coherent deceptive disposition.We trained some mid-sized models on their own plausible but false reasoning.True and false training usually produced nearly identical downstream effects.Even statements contradicting latent knowledge transferred only weakly to unrelated dishonesty.General deception may require agency, persistent private information, and successful concealment over time.IntroductionCurrent frontier models are mundanely misalign...

Read full article →

Related Articles

Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 21h ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago
Why are European countries moving their gold out of North America?
ranit · Hacker News · 13h ago
Can AI design circuit boards yet?
iopapa · Hacker News · 23h ago
Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 1d ago