The Next Web reports a 99.9% vs. 62.7% benchmark gap for the same model and test
The Next Web says OpenAI published the higher score, while the benchmark’s own software returned the lower one. The model and test were the same; the surrounding software setup differed.
TLDR
OpenAI published a 99.9% benchmark score, but the benchmark’s own software returned 62.7%, The Next Web reports. According to the outlet, both results used the same model and test with different supporting software. The Next Web also reports that ARC Prize printed both numbers and says it is not claiming artificial general intelligence.
Combined views
74K
3 Sources, first seen 22d ago