Debate with Self-Play Best-of-N Optimization

·LessWrong··

Context: This is the first research output from Arcadia Alignment’s scalable oversight team, carried out in collaboration with external researchers and mentors (Simon and Jacob). We aim to do rigorous empirical work on debate - bridging the gap from theory to the alignment tasks we care about.Debate is a proposed protocol for scalable oversight. As tasks outrun direct supervision, labs are increasingly likely to train against protocols like it. Our concern is that, for questions which are hard t...

Read full article →

Related Articles

Slovakia finds Russian backdoor in traffic speed cameras
dredmorbius · Hacker News · 5h ago
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
ed-is-ai · Hacker News · 3h ago
Coconut Oil Jet Fuel Matches Kerosene's Efficiency in Engine Tests
mdp2021 · Hacker News · 4h ago
There's no reason for software to be slow anymore
Jach · Hacker News · 1d ago
JIT Compiling Code in 5μs
zX41ZdbW · Hacker News · 14h ago