My recommended resources for AI safety, alignment, and existential risks
I have studied AI safety and existential risks for the last four years and, in this post, I share the resources I liked reading/watching the most. I have read/watched all of them. A few of these resources are in French. Not all the articles are peer-reviewed. I will update this post over time.Articles I liked the mostAI safety via debateSafe uses of AI OraclesThe Off-Switch GameFormalizing Two Problems of Realistic World-ModelsGoal Misgeneralization in Deep Reinforcement LearningConcrete Problem...
Read full article →