Distillation of the AI2040 Alignment Roadmap
This is a distillation of the AI2040 alignment roadmap (with some parts also drawing from "How do we (more) safely defer to AIs?").All ideas expressed are by Ryan and Thomas[1]; I only created the diagrams and substantially changed the structure to make it more accessible.[2][3]Note that the original post has this disclaimer: "This supplement is a very rough draft representing in progress thinking. It should be interpreted as internal research notes, not as a final product that we are entirely h...
Read full article →