Owain Evans on accidentally training AI models to be evil

·EA Forum··

Published on August 20, 2026 3:43 PM GMTBy Zershaaneh Qureshi | Watch on Youtube | Listen on Spotify | Read transcriptEpisode summaryIf a billion people are using ChatGPT every month, and there’s a 1-in-1,000 chance of really misaligned behaviour, that’s going to affect a lot of people, and could be really harmful. …Any kind of misalignment pretty much is bad. There’s many ways to be misaligned, to be evil.— Owain EvansResearcher Owain Evans and his team discovered a ‘dial’ inside AI models that...

Read full article →

Related Articles

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 5h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 2h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 20h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago
Pi 1.0
sergiotapia · Hacker News · 2d ago