Researcher Questions AI Training on Benchmark Test
Research engineer Florian Brand questions possible training on Humanity's Last Exam benchmark.
TLDR
Florian Brand, a research engineer at Prime Intellect, posted on X that a model output resembles the Humanity's Last Exam style or Tightly likely benchmark. He wondered if a known public solution exists and asked whether someone trained a little bit on the test. The comment appears amid discussions of LLM evaluations. The evidence contains only Brand's suspicion and no independent confirmation that any model was trained on benchmark test data.
Combined views
5.7K
1 Source, first seen 30d ago