An "Anthropic Principle" for Formulations of AI Alignment

·LessWrong··

This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I explore one path toward formulating AI alignment without needing to formalize humans or our values, looking instead at the well-integrated computational power characteristic of agents capable of confronting the alignment problem.One natural formulation of AI alignment, the study of how to be sure our highly capable ...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 21h ago
Asahi Linux Now Officially Supports Apple M3 Macs – With Caveats
mdp2021 · Hacker News · 4h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 22h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago