Why Should Corrigible Agents Favor the Present?

·LessWrong··

A corrigible agent understands that it is flawed and seeks to empower its principal to correct those flaws. Many of the intuitive examples of corrigibility happen over a short period of time: the principal gives a command, and then the agent follows it. When there are contradictory commands given at different times, it seems like the agent should follow the most recent command, but it's unclear exactly why corrigible agents favor the present over the past. In this post, I will explore this open ...

Read full article →

Related Articles

Hackers Got Inside a Flock Camera
driverdan · Hacker News · 14h ago
Apple Reference Image: A New Approach for Verified Photography
imwally · Hacker News · 1d ago
Training a 4B model to produce 81% faster query plans than Postgres
polyphilz · Hacker News · 9h ago
Nvidia announces native GPU programming in Rust
nonmaskable · Hacker News · 17h ago
Xiaomi Mimo 2.6 live post-training dashboard
krackers · Hacker News · 8h ago