OpenAI Questions Anthropic ARC-AGI-3 Benchmark Claims
Debate emerges after OpenAI details benchmark settings that altered ARC-AGI-3 results claimed by Anthropic.
TLDR
Anthropic posted a victory tweet stating Claude Opus 5 outperformed competitors on the ARC-AGI-3 benchmark for novel problem-solving. OpenAI research engineer Peter Steinberger publicly questioned the numbers as absurd before posting. OpenAI then published details showing two API settings tripled scores by retaining reasoning and enabling compaction. ARC Prize Foundation president Greg Kamradt responded that the organization applies the same standard harness across providers for consistency. The foundation released its open-source benchmarking repository to support uniform evaluation. Critics noted the harness may need updates beyond simple token in and out comparisons.
Combined views
3.2M
10 Sources, first seen 63d ago