• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    AI models compared with expert judgments on 534 suspicious or borderline math contest submissions

    A post about Repovive’s first math contest says its team worked with IOI, IMO and ICPC finalists to evaluate an AI judge.

    S|
    1 Source, 8h ago, first seen 8h ago

    TLDR

    A post about Repovive’s first math contest says its team worked with IOI, IMO and ICPC finalists to evaluate an AI judge. The team selected 534 suspicious or borderline submissions for review, set a shared grading standard and created expert reference judgments. It then compared models under different grading instructions against those judgments. The post says accuracy is higher across the full set of contest submissions.

    Combined views

    34.9K

    1 Source, first seen 8h ago

    Combined views

    34.9K

    1 Source, first seen 8h ago

    8 likes
    8 likes
    1 comments
    4 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    1 comments
    4 saves
    5 reposts

    1 Source

    @ShayanCJahan1/4 How well can models judge mathematical reasoning? After our first math contest on Repovive, we worked with a team of IOI, IMO, and ICPC finalists to evaluate our AI judge. We selected 534 suspicious or borderline submissions for review, established a shared grading standard, and created expert reference judgments. We then compared models under different grading instructions against those judgments. The main question: can models reliably understand and check challenging mathematical proofs? Accuracy is higher across the full set of contest submissions.8h

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @ShayanCJahan1/4 How well can models judge mathematical reasoning? After our first math contest on Repovive, we worked with a team of IOI, IMO, and ICPC finalists to evaluate our AI judge. We selected 534 suspicious or borderline submissions for review, established a shared grading standard, and created expert reference judgments. We then compared models under different grading instructions against those judgments. The main question: can models reliably understand and check challenging mathematical proofs? Accuracy is higher across the full set of contest submissions.8h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet