Safety's Second Way

·LessWrong··

Epistemics: I've tried to strike a balance between getting it right and getting it out while the community is discussing how to update. I am using the Hack as an example of a broader problem. I look forward to counterarguments.The OpenAI hacks [1] demonstrate an overweighting on single-agent risks, both at OpenAI and within the broader LessWrong and Alignment Forum safety communities. OpenAI neglected to monitor known multi-agent risks even after observing them on their deployment, showing a lac...

Read full article →

Related Articles

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 15h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 14h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 22h ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 17h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 14h ago