Debate with Self-Play Best-of-N Optimization

·LessWrong··

Context: This is the first research output from Arcadia Alignment’s scalable oversight team, carried out in collaboration with external researchers and mentors (Simon and Jacob). We aim to do rigorous empirical work on debate - bridging the gap from theory to the alignment tasks we care about.Debate is a proposed protocol for scalable oversight. As tasks outrun direct supervision, labs are increasingly likely to train against protocols like it. Our concern is that, for questions which are hard t...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 1d ago
The Mathocalypse
6bitquant · Hacker News · 6h ago
Shipping JPEG XL in Chrome
AshleysBrain · Hacker News · 14h ago
Navier–Stokes Lost in Translation
nill0 · Hacker News · 10h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 1d ago