Runtime Assurance Meets Corrigibility: What Happens When an Independent Safety Boundary Blocks an AI Agent? by Idorenyin
This is a research proposal and request for criticism. The experiments and benchmark results are not being reported in this post.The research questionHow can a local safety layer stop an AI agent from operating machinery when doing so would put a nearby human at risk? Imagine an AI agent interacts with a physical environment such as opening a high pressure valve or starting a motor while someone is in harm’s way, the agent requests an action but the safety lay...
Read full article →