How well does AI peer review work?

·Marginal Revolution··

Claude and I planted 100 known errors into 10 open-access psychology papers and then ran them through frontier models and two commercial AI review tools. In brief: The best single system caught 71 of 100 errors, while the worst caught 30. Pooling every system’s output caught 93 of 100. Models are only partly correlated in the errors they find, making ensembling a big lever for finding issues in papers. Check your papers against multiple models! Seven errors could not be caught by any system. All...

Read full article →

Related Articles

UK Fuel Price Intelligence – Market analytics from reporting stations
theazureguy · Hacker News · 3mo ago
Microsoft filings suggest "around 70%" of its AI revenue is on OpenAI
speckx · Hacker News · 4d ago
The iPhone explains 33–52% of fertility decline among women aged 15–44
delichon · Hacker News · 2mo ago
The economics of H-1B immigration
Tyler Cowen · Marginal Revolution · 7d ago
How does the market regard stablecoins?
Tyler Cowen · Marginal Revolution · 7d ago