How well does AI peer review work?
Claude and I planted 100 known errors into 10 open-access psychology papers and then ran them through frontier models and two commercial AI review tools. In brief: The best single system caught 71 of 100 errors, while the worst caught 30. Pooling every system’s output caught 93 of 100. Models are only partly correlated in the errors they find, making ensembling a big lever for finding issues in papers. Check your papers against multiple models! Seven errors could not be caught by any system. All...
Read full article →