• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Teortaxes Claims AI Ability Benchmarks Have Lost Value

    AI researchers discuss the relevance of current model evaluation approaches on X.

    ('
    GM
    T(
    9 Sources, 27d ago, first seen 27d ago

    TLDR

    Teortaxes, identified as a pseudonymous AI developer and commentator known for boosting DeepSeek, stated that the field is generally finished with benchmarks focused on the ability to do X. According to the post, labs can target the next announced hill and surpass it rapidly. Everything reduces to environment, the message continues, leaving only EdgeBench-type meta evaluations. Professor Yoav Goldberg retweeted the statement from his account. Herbie Bradley, a researcher, replied directly to the retweet inquiring about human in the loop or centaur evaluations. This exchange unfolded in the conversation around the original post on X.

    Combined views

    53.8K

    9 Sources, first seen 27d ago

    Combined views

    53.8K

    9 Sources, first seen 27d ago

    468 likes
    468 likes
    64 comments
    98 saves
    33 reposts
    64 comments
    98 saves
    33 reposts

    Sentiment

    Positive24.3%75.7%Negative

    Summary

    Sentiment

    Positive24.3%75.7%Negative

    Replies questioned whether saturating benchmarks like ARC signals real intelligence or AGI, arguing instead that such tests ignore trust, real-world judgment, credentials, and any hallucination risk.

    Based on 45 sentiment-bearing replies from 37 accounts across 2 conversations.

    Summary

    Replies questioned whether saturating benchmarks like ARC signals real intelligence or AGI, arguing instead that such tests ignore trust, real-world judgment, credentials, and any hallucination risk.

    Based on 45 sentiment-bearing replies from 37 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    9 Sources

    @burkovJust because someone puts the word AGI in their benchmark doesn't make it even remotely a reflection of what the general public understands by this term. As a layman, I define AGI as a virtual character that I can reliably ask to do some work a reasonable human with the adequate level of expertise can do, and I can expect that this work will be done the way I described it and reasonable clarfication questions will be asked if needed. Ask yourself a question: would you delegate this AI to negotiate and sign a multi-thousand dollar contract on your behalf, representing your interests with an adequate level of qualification? If the answer is no, it's not AGI.
    @teortaxesTexI think we're generally done with "ability to do X" benchmarks. Labs can just point their machines at the next hill the moment it's announced, and in a few months it's not just conquered but smashed to bits. Everything is environment. EdgeBench-type meta evals are all we have now
    @GaryMarcusRT @burkov: Just because someone puts the word AGI in their benchmark doesn't make it even remotely a reflection of what the general public…
    @Swarooprm7With benchmarks saturating, what’s a true measure of intelligence now? Self promotions of papers/eval strategies welcome too :)
    @shyamalanadkatfeeling the agi?
    @yoavgoRT @teortaxesTex: I think we're generally done with "ability to do X" benchmarks. Labs can just point their machines at the next hill the m…
    @herbiebradley@yoavgo human in the loop/centaur evals?
    @PMinervinihidden test sets! I've received a ton of requests for even the data used in this quick blog post (link in thread) -- I think we're overfitting on test sets

    9 Sources

    @burkovJust because someone puts the word AGI in their benchmark doesn't make it even remotely a reflection of what the general public understands by this term. As a layman, I define AGI as a virtual character that I can reliably ask to do some work a reasonable human with the adequate level of expertise can do, and I can expect that this work will be done the way I described it and reasonable clarfication questions will be asked if needed. Ask yourself a question: would you delegate this AI to negotiate and sign a multi-thousand dollar contract on your behalf, representing your interests with an adequate level of qualification? If the answer is no, it's not AGI.
    @teortaxesTexI think we're generally done with "ability to do X" benchmarks. Labs can just point their machines at the next hill the moment it's announced, and in a few months it's not just conquered but smashed to bits. Everything is environment. EdgeBench-type meta evals are all we have now
    @GaryMarcusRT @burkov: Just because someone puts the word AGI in their benchmark doesn't make it even remotely a reflection of what the general public…
    @Swarooprm7With benchmarks saturating, what’s a true measure of intelligence now? Self promotions of papers/eval strategies welcome too :)
    @shyamalanadkatfeeling the agi?
    @yoavgoRT @teortaxesTex: I think we're generally done with "ability to do X" benchmarks. Labs can just point their machines at the next hill the m…
    @herbiebradley@yoavgo human in the loop/centaur evals?
    @PMinervinihidden test sets! I've received a ton of requests for even the data used in this quick blog post (link in thread) -- I think we're overfitting on test sets