We cannot safely automate value alignment evaluation and research without thinking about delegation and discretion by Maria Federica Martino Lena

·Nuno Sempere··

This es­s­say was origi­nally posted on LessWrong at https://​​www.less­wrong.com/​​posts/​​JrRWGQKjCiLRYXSjW/​​we-can­not-safely-au­to­mate-value-al­ign­ment-eval­u­a­tion-and-1IntroductionAu­tomat­ing al­ign­ment eval­u­a­tion and re­search is thought to be an effi­cient way to safe­guard against un­con­trol­lable AGI, as Joe Car­l­smith and Jan Leike them­selves ad­mit­ted. In par­tic­u­lar, Leike pro­posed a Min­i­mal Vi­able Product (MVP) for al­ign­ment, con­sist­ing in:“Build­ing a suffi­...

Read full article →

Related Articles

34% of the US public is now aware of AI xrisk, and the curve is steepening
otto.barten · LessWrong · 42m ago
34% of the US public is now aware of AI xrisk, and the curve is steepening by Otto
Otto · Nuno Sempere · 40m ago
New Research: Climate Mitigation is Overlooked by EA by Dan Stein
Dan Stein · Nuno Sempere · 1d ago
Will any non-astronaut be "commuting to the moon" before 2040?
Panfilo · Manifold Markets · 1d ago
Lexical Filtering: Deciding under Indeterminacy with Lexically Ordered Values by Kuutti Lappalainen
Kuutti Lappalainen · Nuno Sempere · 2d ago