OpenAI Still Benchmarks Frontier Models on GPQA Diamond
Researcher notes continued use of an established AI evaluation test.
TLDR
Djamé Seddah, associate professor and NLP researcher at INRIA Paris, retweeted a post from Benjamin Marie. The post states that OpenAI continues benchmarking its frontier models on GPQA Diamond. It adds that the benchmark now ranks among the oldest still in common use. The comment draws attention to the persistence of this particular evaluation set even as newer tests appear. No further details on specific model scores or timelines appear in the shared post. The retweet brings the remark to an academic audience focused on computational linguistics and multilingual work.
Combined views
1
1 Source, first seen 27d ago