AI's progress on hard math and the limits of benchmarks
TWIML AI's episode with Greg Burnham, who leads AI capabilities research at Epoch AI, examines how models tackle difficult math problems and whether they're producing genuinely new ideas.
TLDR
TWIML AI says AI systems are helping tackle long-standing research problems, including Navier-Stokes. Its episode with Burnham explores how much models rely on persistence and prior human work, how to measure progress as traditional benchmarks become less useful, and where models still struggle with open-ended work, learning from experience and identifying promising research directions.
Combined views
—
1 Source, first seen 12h ago
