Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)

·LessWrong··

1.1 Tl;drAlignment is often conceptualized as AIs helping humans achieve their goals: AIs that increase people’s agency and empowerment; AIs that are helpful, corrigible, and/or obedient; AIs that avoid manipulating people. But that last one—manipulation—points to a challenge for all these desiderata: a human’s goals are themselves under-determined and manipulable, and it’s awfully hard to pin down a principled distinction between changing people’s goals in a good way (“providing counsel”, “prov...

Read full article →

Related Articles

Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 14h ago
Moderna reports first positive Phase 3 for mRNA neoantigen therapy in melanoma
heydenberk · Hacker News · 12h ago
Google replaced Git tags for certain source code with obtaining via Google Drive
Animux · Hacker News · 8h ago
Go 1.27
database64128 · Hacker News · 7h ago
Mathematics in the age of AI
jonbaer · Hacker News · 11h ago