How's it going? Reinforcement learning in language models recruits a functional welfare axis

·LessWrong··

In collaboration with David Chalmers and Pavel Izmailov. Work done at NYU. Andy wrote this summary of the paper, which you can find in full on the website, or, if you insist on a PDF, arXiv.IntroductionWe know that language models work in a vast and shadowy landscape of entanglements and associations. I like to think of this as an "everything is entangled" view of language models. Emergent misalignment fits, indeed helped define, this frame. If you reward bad stuff, then the model gets generally...

Read full article →

Related Articles

New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 6h ago
A walk through of the DeltaNet family of linear attention variants
AnhTho_FR · Hacker News · 3h ago
GrapheneOS Defends Data-Wiping Function That Blocked US Border Search
pseudolus · Hacker News · 4h ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 1d ago
Zig's Incremental Compilation Internals
garyhtou · Hacker News · 4h ago