Teortaxes Claims AI Ability Benchmarks Have Lost Value
AI researchers discuss the relevance of current model evaluation approaches on X.
TLDR
Teortaxes, identified as a pseudonymous AI developer and commentator known for boosting DeepSeek, stated that the field is generally finished with benchmarks focused on the ability to do X. According to the post, labs can target the next announced hill and surpass it rapidly. Everything reduces to environment, the message continues, leaving only EdgeBench-type meta evaluations. Professor Yoav Goldberg retweeted the statement from his account. Herbie Bradley, a researcher, replied directly to the retweet inquiring about human in the loop or centaur evaluations. This exchange unfolded in the conversation around the original post on X.
Combined views
53.8K
9 Sources, first seen 27d ago