When capabilities work is the *safe* bet

·LessWrong··

If you believe that LLMs lend themselves unusually well to alignment compared to other regimes, this can be a very good reason to start doing capability research on them rather than LLM safety research. Imagine you have these beliefs about how AI goes: mjx-container[jax="CHTML"] { line-height: 0; } mjx-container [space="1"] { margin-left: .111em; } mjx-container [space="2"] { margin-left: .167em; } mjx-container [space="3"] { margin-left: .222em; } mjx-container [space="4"] { margin-left: .278em...

Read full article →

Related Articles

Livenerf: Has Opus 5.5 been nerfed yet?
bryan0 · Hacker News · 14h ago
How Delhi cut electricity loss from 50 to 5 percent
rbanffy · Hacker News · 1d ago
Vermont replacing power plants with home batteries
devonnull · Hacker News · 19h ago
US sanctions force The Netherlands off Microsoft and toward alternative NixOS
mywacaday · Hacker News · 1d ago
500k facial scans at UK stations yield no arrests, 1 false positive
ilamont · Hacker News · 1d ago