• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    BrokenArXiv and ArXivMath updated to focus on recently refuted conjectures

    The release announcement puts GPT-6 Astra on top and says models now run in a test harness rather than directly through an API.

    TiboTI
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)T(
    Vitaliy ChileyVC
    6 Sources, ,

    TLDR

    The September 16, 2026, release announcement for BrokenArXiv and ArXivMath says the updated benchmarks focus on conjectures refuted on arXiv in the preceding month. It describes a shift to running models in a test harness instead of directly via API, and reports GPT-6 Astra on top.

    Combined views

    3.8M

    6 Sources, first seen 21d ago

    Combined views

    3.8M

    6 Sources, first seen 21d ago

    7.7K likes
    21d ago
    first seen 21d ago
    7.7K likes
    1.7K comments
    1.2K saves
    380 reposts
    1.7K comments
    1.2K saves
    380 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    Jasper Dekoninck@j_dekoninckWe are releasing the latest version of BrokenArXiv and ArXivMath! These benchmarks now focus on conjectures that were refuted in the last month on ArXiv, and models are executed within a harness instead of directly via API. Performance remains impressive, with GPT-6 Astra on top21d
    Vitaliy Chiley@vitaliychileyRT @j_dekoninck: We are releasing the latest version of BrokenArXiv and ArXivMath! These benchmarks now focus on conjectures that were refu…21d
    Tibo@thsottiauxAstra ✅ Fast ✅ Frontier ✅ Efficient ✅ For everyone21d
    Brian Roemmele@BrianRoemmeleTake a ganders at the most expensive AI model and how it is rapidly dropping in performance as prices are outrageous. The is the BrokenArXiv and ArXivMath benchmarks with focus on conjectures that were refuted in the last month on ArXiv, and models are executed within a harness instead of directly via API. This poor showing may explain the fear and doom theater from the Antropic cult.21d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexRT @j_dekoninck: Unfortunately, we had to remove this entry: the authors ran this as an autoformalization task with access to source papers…20d

    6 Sources

    Jasper Dekoninck@j_dekoninckWe are releasing the latest version of BrokenArXiv and ArXivMath! These benchmarks now focus on conjectures that were refuted in the last month on ArXiv, and models are executed within a harness instead of directly via API. Performance remains impressive, with GPT-6 Astra on top21d
    Vitaliy Chiley@vitaliychileyRT @j_dekoninck: We are releasing the latest version of BrokenArXiv and ArXivMath! These benchmarks now focus on conjectures that were refu…21d
    Tibo@thsottiauxAstra ✅ Fast ✅ Frontier ✅ Efficient ✅ For everyone21d
    Brian Roemmele@BrianRoemmeleTake a ganders at the most expensive AI model and how it is rapidly dropping in performance as prices are outrageous. The is the BrokenArXiv and ArXivMath benchmarks with focus on conjectures that were refuted in the last month on ArXiv, and models are executed within a harness instead of directly via API. This poor showing may explain the fear and doom theater from the Antropic cult.21d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexRT @j_dekoninck: Unfortunately, we had to remove this entry: the authors ran this as an autoformalization task with access to source papers…20d