Owain Evans on accidentally training AI models to be evil
Published on August 20, 2026 3:43 PM GMTBy Zershaaneh Qureshi | Watch on Youtube | Listen on Spotify | Read transcriptEpisode summaryIf a billion people are using ChatGPT every month, and there’s a 1-in-1,000 chance of really misaligned behaviour, that’s going to affect a lot of people, and could be really harmful. …Any kind of misalignment pretty much is bad. There’s many ways to be misaligned, to be evil.— Owain EvansResearcher Owain Evans and his team discovered a ‘dial’ inside AI models that...
Read full article →