The best system caught 71 errors out of 100 — This indicates significant room for improvement in current AI tools, suggesting they are not yet fully reliable.
Seven unique omissions were missed by all systems — Highlighting the need to be cautious and possibly double-check critical information manually.
Further testing with more advanced models could provide better insights.
Claude and I planted 100 known errors into 10 open-access psychology papers and then ran them through frontier models and two commercial AI review tools. In brief: The best single system caught 71 of 100 errors, while the worst caught 30.