Can Frontier Models Autocomplete Safety Research?

·LessWrong··

We pose the following research question: how can we measure the "research taste" of language models in experiment planning? What parts of planning taste remain intrinsic to humans? TL;DR. A future where “tasteless autoresearch” improves capabilities but not safety is plausible and dangerous. We need rough tests of tasteful planning to see what is missing.We can start by masking part of a paper, sampling extensions from a language model, and comparing them against the masked experiments. This pro...

Read full article →

Related Articles

Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
garo-pro · Hacker News · 8h ago
FDA authorizes first wearable device that monitors ketone and blood sugar levels
sunnynagra · Hacker News · 23h ago
Xiaomi: New CPU matches Apple cores single threaded, much faster multithreaded
tosh · Hacker News · 2d ago
Black hole singularity is a surface not a point
raattgift · Hacker News · 1d ago
GLM-5.3-Flash Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 3h ago