Quantilized debate and consultancy in image environments: protocol design lessons for scalable oversight experiments

·LessWrong··

TL;DRWe present an improved version of the sparse-pixel experiment from AI Safety via Debate (2018), where a weak judge classifies an image from 4 of its 36 cells and powerful agents (in our case, exact minimax against the judge) choose which cells the judge sees. Specifically, an agent called mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse...

Read full article →

Related Articles

Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 6h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 19h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 13h ago