OpenAI published 99.9%; benchmark software returned 62.7%, The Next Web reports
The Next Web says the scores involved the same model and test with different surrounding software setups. The outlet reports that ARC Prize published both figures and says it is not claiming artificial general intelligence.
TLDR
The Next Web reports that OpenAI published a 99.9% benchmark result, while the benchmark’s own software returned 62.7%. It describes the same model and test with different surrounding software setups, and says ARC Prize printed both numbers without claiming artificial general intelligence (AGI). A critic sharing the report accused OpenAI of pretending Astra was AGI and called for the company to be shut down until there are “changes at the top.”
Combined views
63.6K
1 Source, first seen 23d ago