We cannot safely automate value alignment evaluation and research without thinking about delegation and discretion by Maria Federica Martino Lena

·Nuno Sempere··

This es­s­say was origi­nally posted on LessWrong at https://​​www.less­wrong.com/​​posts/​​JrRWGQKjCiLRYXSjW/​​we-can­not-safely-au­to­mate-value-al­ign­ment-eval­u­a­tion-and-1IntroductionAu­tomat­ing al­ign­ment eval­u­a­tion and re­search is thought to be an effi­cient way to safe­guard against un­con­trol­lable AGI, as Joe Car­l­smith and Jan Leike them­selves ad­mit­ted. In par­tic­u­lar, Leike pro­posed a Min­i­mal Vi­able Product (MVP) for al­ign­ment, con­sist­ing in:“Build­ing a suffi­...

Read full article →

Related Articles

What is recursive self-improvement, and what would it mean to ban it? by sarahhw
sarahhw · Nuno Sempere · 22h ago
State of the Field: AI for Epistemics and Coordination by Ben_N
Ben_N · Nuno Sempere · 22h ago
Will a non-lab AI-only effort solve a Millennium Problem in 2026?
Bayesian · Manifold Markets · 1d ago
Will the U.S. use a nuclear weapon against Iran’s Pickaxe Mountain between November 4 and November 10, 2026?
Laurent27bis · Manifold Markets · 1d ago
How many people will attend EAG NYC 2026?
d · Manifold Markets · 1d ago