Announcement
AI models compared with expert judgments on 534 suspicious or borderline math contest submissions
A post about Repovive’s first math contest says its team worked with IOI, IMO and ICPC finalists to evaluate an AI judge.
TLDR
A post about Repovive’s first math contest says its team worked with IOI, IMO and ICPC finalists to evaluate an AI judge. The team selected 534 suspicious or borderline submissions for review, set a shared grading standard and created expert reference judgments. It then compared models under different grading instructions against those judgments. The post says accuracy is higher across the full set of contest submissions.
Combined views
34.9K
1 Source, first seen ago
