AI Alignment at Which Abstraction Level?
This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I start from the observation that AI alignment usually takes for granted a privileged role for human values, without asking why values should come from that particular abstraction level. I then give an independent engineering reason to worry about putting humans at the center of an alignment specification, drawing on ...
Read full article →