Training a Conceptual Reasoning Judge

·LessWrong··

TL;DR: We fine-tune a judge LLM on our conceptual reasoning dataset to output a critique rating in a single forward pass. This method provides significant uplift in performance on held-out critiques, measured by alignment with our human expert ratings. We find that improvement is mostly concentrated in discriminating low-quality, often model-written critiques. However, our trained model still retains its improvement over the base model on rewrites of our critiques, meaning that improvement canno...

Read full article →

Related Articles

GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 17h ago
In Australia, a home battery boom has helped cut wholesale power prices
speckx · Hacker News · 9h ago
Going Dark, and the era of law enforcement hacking
vslira · Hacker News · 2h ago
Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 4h ago
Where did the old web go? We followed 657,607 links to find out
tdx · Hacker News · 1d ago