Training AI to be better at correctness than persuasion

·LessWrong··

Cross-posted from my website. I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. "Super-persuasive" AI is dangerous because a misaligned ASI could persuade humans to help it take over. But setting that aside, even if we manage to make ASI friendly, it may provide super-persuasive but mistaken guidance that permanently sets us down the wrong path. This post focuses on the danger of an aligned ...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 8h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 11h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 1d ago
Turns are Better than Radians (2022)
mayoff · Hacker News · 20h ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago