Owain Evans on accidentally training AI models to be evil

·EA Forum··

Published on August 20, 2026 3:43 PM GMTBy Zershaaneh Qureshi | Watch on Youtube | Listen on Spotify | Read transcriptEpisode summaryIf a billion people are using ChatGPT every month, and there’s a 1-in-1,000 chance of really misaligned behaviour, that’s going to affect a lot of people, and could be really harmful. …Any kind of misalignment pretty much is bad. There’s many ways to be misaligned, to be evil.— Owain EvansResearcher Owain Evans and his team discovered a ‘dial’ inside AI models that...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 4h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 7h ago
Turns are Better than Radians (2022)
mayoff · Hacker News · 16h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 23h ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago