When capabilities work is the *safe* bet

·LessWrong··

If you believe that LLMs lend themselves unusually well to alignment compared to other regimes, this can be a very good reason to start doing capability research on them rather than LLM safety research. Imagine you have these beliefs about how AI goes: mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em...

Read full article →

Related Articles

AI has access to a vastly larger working memory than the human brain
rzk · Hacker News · 14h ago
Semaglutide linked to lower predicted dementia risk
randycupertino · Hacker News · 16h ago
Abdominal fat predicts heart disease risk better than BMI
theanonymousone · Hacker News · 11h ago
Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 1d ago
What happens when an LLM never sees material beyond fifth grade?
porridgeraisin · Hacker News · 1h ago