• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    The Next Web reports a 99.9% vs. 62.7% benchmark gap for the same model and test

    The Next Web says OpenAI published the higher score, while the benchmark’s own software returned the lower one. The model and test were the same; the surrounding software setup differed.

    GM
    DJ
    BM
    3 Sources, ,

    TLDR

    OpenAI published a 99.9% benchmark score, but the benchmark’s own software returned 62.7%, The Next Web reports. According to the outlet, both results used the same model and test with different supporting software. The Next Web also reports that ARC Prize printed both numbers and says it is not claiming artificial general intelligence.

    Combined views

    74K

    3 Sources, first seen 22d ago

    Combined views

    74K

    3 Sources, first seen 22d ago

    512 likes
    22d ago
    first seen 22d ago
    512 likes
    67 comments
    169 saves
    43 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    67 comments
    169 saves
    43 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @GaryMarcusso it was the harness not the model. more sneaky games with unacknowledged domain-specific neurosymbolic AI doing the real work.
    @bnjmn_marieOne of the most interesting tables in the DeepSeek V4.1 model card. Same model, same benchmark, wildly different scores depending on the agent harness. One more proof that comparing your new model’s score with previously published numbers is close to meaningless unless the harness and config are matched.
    @zehavocRT @bnjmn_marie: One of the most interesting tables in the DeepSeek V4.1 model card. Same model, same benchmark, wildly different scores d…

    3 Sources

    @GaryMarcusso it was the harness not the model. more sneaky games with unacknowledged domain-specific neurosymbolic AI doing the real work.
    @bnjmn_marieOne of the most interesting tables in the DeepSeek V4.1 model card. Same model, same benchmark, wildly different scores depending on the agent harness. One more proof that comparing your new model’s score with previously published numbers is close to meaningless unless the harness and config are matched.
    @zehavocRT @bnjmn_marie: One of the most interesting tables in the DeepSeek V4.1 model card. Same model, same benchmark, wildly different scores d…