An "Anthropic Principle" for Formulations of AI Alignment
This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I explore one path toward formulating AI alignment without needing to formalize humans or our values, looking instead at the well-integrated computational power characteristic of agents capable of confronting the alignment problem.One natural formulation of AI alignment, the study of how to be sure our highly capable ...
Read full article →