Constitutional AI Alignment

·LessWrong··

Epistemological status: existential AI safety suggestions as a set of bullet pointsTL;DR We need to be clear about what behavior we’re trying to train when we align AIs — aligned behavior is not the default for an LLM whose behavior is distilled from ours, nor for an RL-trained maximizer. Anthropic published an early constitution, and recently a large update in Claude’s Constitution (a.k.a. “Soul Doc”). That includes not just what the AI should do, but explanations of how and why many decisions ...

Read full article →

Related Articles

New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 7h ago
A walk through of the DeltaNet family of linear attention variants
AnhTho_FR · Hacker News · 4h ago
GrapheneOS Defends Data-Wiping Function That Blocked US Border Search
pseudolus · Hacker News · 5h ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 1d ago
Zig's Incremental Compilation Internals
garyhtou · Hacker News · 4h ago