Training AI to be better at correctness than persuasion

·LessWrong··

Cross-posted from my website. I continue to believe we should pause frontier AI development. Any discussion of alternative strategies should be thought of as planning for contingencies. "Super-persuasive" AI is dangerous because a misaligned ASI could persuade humans to help it take over. But setting that aside, even if we manage to make ASI friendly, it may provide super-persuasive but mistaken guidance that permanently sets us down the wrong path. This post focuses on the danger of an aligned ...

Read full article →

Related Articles

Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 4h ago
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 11h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 8h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 1d ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago