Claude is Now Alignment-Pretrained

·LessWrong··

Anthropic are now actively using the approach to alignment often called “Alignment Pretraining” or “Safety Pretraining” — using Stochastic Gradient Descent on a large body of natural or synthetic documents showing the AI assistant doing the right thing in morally challenging situations. They tried this out, found it works well and generalizes well, and they’re now using it.I’m absolutely delighted. I’ve been repeatedly advocating this approach on LessWrong and the Alignment Forum for a couple of...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 9h ago
Pi 1.0
sergiotapia · Hacker News · 1d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 23h ago
FTL: A new operating system for clouds
romac · Hacker News · 4h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago