Debate with Self-Play Best-of-N Optimization

·LessWrong··

Context: This is the first research output from Arcadia Alignment’s scalable oversight team, carried out in collaboration with external researchers and mentors (Simon and Jacob). We aim to do rigorous empirical work on debate - bridging the gap from theory to the alignment tasks we care about.Debate is a proposed protocol for scalable oversight. As tasks outrun direct supervision, labs are increasingly likely to train against protocols like it. Our concern is that, for questions which are hard t...

Read full article →

Related Articles

Data centers raise nearby temperatures by up to 4 degrees in Phoenix
cwwc · Hacker News · 3h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 12h ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 8h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
India has paved the way for charging merchants a fee on UPI transactions
monkey_monkey · Hacker News · 1d ago