Training a Conceptual Reasoning Judge

·LessWrong··

TL;DR: We fine-tune a judge LLM on our conceptual reasoning dataset to output a critique rating in a single forward pass. This method provides significant uplift in performance on held-out critiques, measured by alignment with our human expert ratings. We find that improvement is mostly concentrated in discriminating low-quality, often model-written critiques. However, our trained model still retains its improvement over the base model on rewrites of our critiques, meaning that improvement canno...

Read full article →

Related Articles

U.S. Strategic Petroleum Reserve Falls to Lowest Level Since 1982
thelastgallon · Hacker News · 7h ago
Does Reddit have an astroturfing problem? What the data suggests
p-s-v · Hacker News · 20h ago
Nvidia wants to put a watchdog chip next to every AI agent
jonbaer · Hacker News · 17h ago
Nissan's third generation e-POWER powertrain
mroche · Hacker News · 1d ago
MicroLLM Lab – Try 7 tiny LLM's in the browser
logicallee · Hacker News · 14h ago