Benjamin Marie Questions GPQA Diamond Benchmark Usage by OpenAI
Independent researcher notes persistence of older evaluation method for advanced AI models.
TLDR
Benjamin Marie, an independent research scientist specializing in NLP and LLMs, posted that OpenAI is still benchmarking frontier models on GPQA Diamond. He described it as one of the oldest benchmarks still commonly used for evaluating top-tier models. Marie questioned why this continues, pointing out that countless variants of the questions are likely now available online in paraphrased, slightly modified, or translated forms.
Combined views
4.5K
1 Source, first seen 27d ago