Why Should Corrigible Agents Favor the Present?

·LessWrong··

A corrigible agent understands that it is flawed and seeks to empower its principal to correct those flaws. Many of the intuitive examples of corrigibility happen over a short period of time: the principal gives a command, and then the agent follows it. When there are contradictory commands given at different times, it seems like the agent should follow the most recent command, but it's unclear exactly why corrigible agents favor the present over the past. In this post, I will explore this open ...

Read full article →

Related Articles

Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 12h ago
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 4h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 4h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 3h ago
Nvidia projects $673B in sales as AI demand widens
kuuuzya · Hacker News · 6h ago