GPT-6 Astra Sets ARC-AGI-3 Record
François Chollet and ARC Prize discuss GPT-6 Astra's benchmark results and what they mean for AGI claims.
TLDR
ARC Prize posted that OpenAI's GPT-6 Astra scores 63 percent on ARC-AGI-3 and reaches 99 percent with a new harness. The system surpasses human performance on 96 percent of tasks and builds precise symbolic models of novel environments. Matt Mazur reported 62.7 percent using their standard harness, more than doubling the prior verified score. François Chollet stated that benchmarking co-evolves with models and that saturating ARC-AGI-3 does not prove AGI. Mike Knoop called the result a large leap toward AGI but noted open-ended invention remains unsolved. Chollet added that progress arrived twice as fast as his earlier one-year estimate.
Combined views
4.6M
68 Sources, first seen 27d ago