AI Reviews Catch More Issues Than Human NeurIPS Reviews
Area chairs claim LLMs produce far superior peer review feedback on papers.
TLDR
Researchers including NeurIPS area chairs report that well-executed LLM reviews identify substantially more flaws than typical human reviews across optimization and machine learning submissions. Discussions note AI strengths in exhaustive consistency and coverage checks while questioning whether higher issue counts improve final decisions. Some participants suggest posting system prompts for uniform standards or replacing most human reviews, yet others highlight differing review objectives and risks of self-preference bias in automated output.
Combined views
77.4K
12 Sources, first seen 68d ago