Leading AI models reportedly nearly max out corrected physics benchmarks
A preprint's author says Yale physicists helped re-grade supposed model failures, finding that the benchmarks—not the models—were often at fault.
TLDR
Announcing the preprint on September 14, 2026, one of its authors said expert re-grading with Yale physicists brought leading models close to maxing out the physics benchmarks studied, including Humanity Last Exam's physics component and CritPt. A separate post summarizing the paper says only 12 of 250 rejected cases across four audited benchmark subsets were actual model mistakes. The other 238 involved bad questions, wrong reference answers or graders rejecting correct answers. That summary also notes an important limit: the authors' agents failed to fully solve any of the open theoretical-physics problems they tried.
Combined views
369.2K
10 Sources, first seen 3d ago
Leading AI models reportedly nearly max out corrected physics benchmarks
A preprint's author says Yale physicists helped re-grade supposed model failures, finding that the benchmarks—not the models—were often at fault.