Userland Alignment

·LessWrong··

Most discourse around AI alignment centers on model development and the labs that develop them. This is a reasonable place to focus given the centrality of model training to AI advancement. However, there are neglected opportunities to build defense-in-depth via aligned harnesses – and these opportunities might be tractable by interested developers and researchers who otherwise would struggle to have impact given the limited opportunities to influence lab practices.The behavior of an AI system i...

Read full article →

Related Articles

Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 6h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 19h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 13h ago