• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Leading AI models reportedly nearly max out corrected physics benchmarks

    A preprint's author says Yale physicists helped re-grade supposed model failures, finding that the benchmarks—not the models—were often at fault.

    Ethan MollickEM
    Ted SandersTS
    Arman CohanAC
    10 Sources, ,

    TLDR

    Announcing the preprint on September 14, 2026, one of its authors said expert re-grading with Yale physicists brought leading models close to maxing out the physics benchmarks studied, including Humanity Last Exam's physics component and CritPt.

    A separate post summarizing the paper says only 12 of 250 rejected cases across four audited benchmark subsets were actual model mistakes. The other 238 involved bad questions, wrong reference answers or graders rejecting correct answers. That summary also notes an important limit: the authors' agents failed to fully solve any of the open theoretical-physics problems they tried.

    Combined views

    369.4K

    10 Sources, first seen 23d ago

    Combined views

    369.4K

    10 Sources, first seen 23d ago

    2.4K likes
    23d ago
    first seen 23d ago
    2.4K likes
    123 comments
    591 saves
    280 reposts
    123 comments
    591 saves
    280 reposts

    Sentiment

    Positive47.8%52.2%Negative

    Summary

    Positive accounts praised the new paper for exposing broken grading and near-saturation in frontier physics benchmarks, while negative replies criticized the evaluations as marketing tools full of errors.

    Based on 48 sentiment-bearing replies from 46 accounts across 4 conversations.

    Sentiment

    Positive47.8%52.2%Negative

    Summary

    Positive accounts praised the new paper for exposing broken grading and near-saturation in frontier physics benchmarks, while negative replies criticized the evaluations as marketing tools full of errors.

    Based on 48 sentiment-bearing replies from 46 accounts across 4 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    10 Sources

    Zeyi(Andy) Liu@ZeyiAndyLiuNew paper: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks Frontier models have tackled Navier–Stokes. People naturally wonder, how is their capability for physical sciences? According to today’s benchmarks, frontier models still struggle with physics (e.g. 47.3% with GPT5.6-sol on Humanity Last Exam - Physics component). But that picture did not match physicists' experience using these models. So we did a rigorous analysis of current benchmarks and re-graded their supposed failures with Yale physicists. We found that the problem was often the benchmark and not the model !! When corrected, the models almost saturate all benchmarks, including everyone’s favourite, Humanity Last Exam (physics-part) and Critpt. These failures in turn make AA less reliable as a measure of model’s true capability. For more details, results and analysis check out our preprint: https://arxiv.org/abs/2609.13009 and blog post: https://jsous.github.io/blogs/is-physics-dead/23d
    Arman Cohan@armancohanRT @ZeyiAndyLiu: New paper: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Le…23d
    Lisan al Gaib@scaling01turns out most benchmarks are trash "When corrected, the models almost saturate all benchmarks, including everyone’s favourite, Humanity Last Exam (physics-part) and CritPt."22d
    Chris@ChrisGPTEver since Artificial Analysis pinned GPT-6 Astra below many others we all felt something was wrong, now we may have evidence. A new paper had Yale physicists manually re-grade the supposed failures of frontier models across major physics benchmarks, and found that a huge amount of the problem was actually the benchmark/grading itself, not the models. This is something @tszzl and many other lab employees note. Once the broken questions and grading were corrected, frontier models nearly saturated several of these benchmarks. GPT-5.6 Sol for example jumps from 47.3% to 78.7% on HLE-Physics mean@4, and hits 94.4% pass@4 on corrected CritPt. This is a pretty brutal indictment of the index and I hope they are able to look into these results!22d
    Ethan Mollick@emollickThe state of public AI benchmarking is dire and is undermining our ability to understand how good AI is now. Most famous measures are maxed out, and, as this paper shows, the non-saturated benchmarks are riddled with so many errors that they vastly underestimate AI abilities.21d
    Ted Sanders@sanderstedRT @ChrisGPT: Ever since Artificial Analysis pinned GPT-6 Astra below many others we all felt something was wrong, now we may have evidence…21d
    Rohan Paul@rohanpaul_aiNew Yale Univ + other top lab paper shows frontier models are already close to maxing out today’s closed-ended physics benchmarks. And many apparent failures come from bad questions, wrong reference answers, and brittle graders explain most audited physics failures, making benchmark quality the new bottleneck for measuring frontier models. The researchers had physicists re-check model failures across 6 popular physics benchmarks instead of trusting the original scores. In 250 rejected cases from 4 audited benchmark subsets, only 12 were actual model mistakes. The other 238 came from bad questions, wrong reference answers, or graders rejecting correct answers. After expert review, GPT-5.6-Sol’s measured HLE-Physics score rose from 47.3% to 78.7%. So a low physics benchmark score can badly underestimate what a frontier model can actually solve. But that does not mean these models can reliably do physics research: the authors’ agents still failed to fully solve any of the open theoretical-physics problems they tried. but it does mean that do not treat benchmark scores as clean ground truth anymore for AI's Physics capability.21d

    10 Sources

    Zeyi(Andy) Liu@ZeyiAndyLiuNew paper: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks Frontier models have tackled Navier–Stokes. People naturally wonder, how is their capability for physical sciences? According to today’s benchmarks, frontier models still struggle with physics (e.g. 47.3% with GPT5.6-sol on Humanity Last Exam - Physics component). But that picture did not match physicists' experience using these models. So we did a rigorous analysis of current benchmarks and re-graded their supposed failures with Yale physicists. We found that the problem was often the benchmark and not the model !! When corrected, the models almost saturate all benchmarks, including everyone’s favourite, Humanity Last Exam (physics-part) and Critpt. These failures in turn make AA less reliable as a measure of model’s true capability. For more details, results and analysis check out our preprint: https://arxiv.org/abs/2609.13009 and blog post: https://jsous.github.io/blogs/is-physics-dead/23d
    Arman Cohan@armancohanRT @ZeyiAndyLiu: New paper: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Le…23d
    Lisan al Gaib@scaling01turns out most benchmarks are trash "When corrected, the models almost saturate all benchmarks, including everyone’s favourite, Humanity Last Exam (physics-part) and CritPt."22d
    Chris@ChrisGPTEver since Artificial Analysis pinned GPT-6 Astra below many others we all felt something was wrong, now we may have evidence. A new paper had Yale physicists manually re-grade the supposed failures of frontier models across major physics benchmarks, and found that a huge amount of the problem was actually the benchmark/grading itself, not the models. This is something @tszzl and many other lab employees note. Once the broken questions and grading were corrected, frontier models nearly saturated several of these benchmarks. GPT-5.6 Sol for example jumps from 47.3% to 78.7% on HLE-Physics mean@4, and hits 94.4% pass@4 on corrected CritPt. This is a pretty brutal indictment of the index and I hope they are able to look into these results!22d
    Ethan Mollick@emollickThe state of public AI benchmarking is dire and is undermining our ability to understand how good AI is now. Most famous measures are maxed out, and, as this paper shows, the non-saturated benchmarks are riddled with so many errors that they vastly underestimate AI abilities.21d
    Ted Sanders@sanderstedRT @ChrisGPT: Ever since Artificial Analysis pinned GPT-6 Astra below many others we all felt something was wrong, now we may have evidence…21d
    Rohan Paul@rohanpaul_aiNew Yale Univ + other top lab paper shows frontier models are already close to maxing out today’s closed-ended physics benchmarks. And many apparent failures come from bad questions, wrong reference answers, and brittle graders explain most audited physics failures, making benchmark quality the new bottleneck for measuring frontier models. The researchers had physicists re-check model failures across 6 popular physics benchmarks instead of trusting the original scores. In 250 rejected cases from 4 audited benchmark subsets, only 12 were actual model mistakes. The other 238 came from bad questions, wrong reference answers, or graders rejecting correct answers. After expert review, GPT-5.6-Sol’s measured HLE-Physics score rose from 47.3% to 78.7%. So a low physics benchmark score can badly underestimate what a frontier model can actually solve. But that does not mean these models can reliably do physics research: the authors’ agents still failed to fully solve any of the open theoretical-physics problems they tried. but it does mean that do not treat benchmark scores as clean ground truth anymore for AI's Physics capability.21d