How risky would it be to make powerful AI obey one or a few people?
It seems fairly likely that the first powerful AIs will be instruction-following rather than value-aligned, and will be controlled by a small number of people. So it makes sense to worry what individual people might do with such immense power. Here intuitions diverge and careful analysis is scarce. This post presents a debate between Seth Herd and cousin_it over how risky such a scenario would be. The debate ran under an unusual protocol. First we wrote our initial draft statements and sent them...
Read full article →