When does debate help a weak judge? Evidence from code and logic

·LessWrong··

Authors: Ethan Elasky and Frank Nakasako, Palaestra Research; Naman Goyal, Independent.ArXiv link: [will be here when available] (for now, preprint Google Drive link)Thanks to Coefficient Giving for support and Thinking Machines for API credits; our mentor for guidance along the way; and Julian Michael, Johannes Gasteiger, and Jiaxin Wen, among others, for helpful conversations. What we didThis is a writeup of experiments we ran on debate as a reward-labeling protocol. The basic question: if a w...

Read full article →

Related Articles

Measuring the sloppiness of code
doppp · Hacker News · 14h ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
HuggingFace: Security.txt
yarapavan · Hacker News · 13h ago
Rune is now open source
ernestrc · Hacker News · 12h ago
The Deathray: A simple way for an untrusted site to freeze a Mac
auberonedu · Hacker News · 1d ago