Safety's Second Way

·LessWrong··

Epistemics: I've tried to strike a balance between getting it right and getting it out while the community is discussing how to update. I am using the Hack as an example of a broader problem. I look forward to counterarguments.The OpenAI hacks [1] demonstrate an overweighting on single-agent risks, both at OpenAI and within the broader LessWrong and Alignment Forum safety communities. OpenAI neglected to monitor known multi-agent risks even after observing them on their deployment, showing a lac...

Read full article →

Related Articles

LG smart TVs caught logging audio with screen off and snooping on local devices
chris_overseas · Hacker News · 13h ago
Smartphone makers don't bother to comply with EU repairability requirements
mdp2021 · Hacker News · 9h ago
bzip3
tosh · Hacker News · 7h ago
Asahi Linux on M3
mdp2021 · Hacker News · 1d ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 1d ago