• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Web-search leaderboard launches with tests of 24 AI models

    The creators say they made their internal evaluation method public to compare search performance end to end, with cost and latency alongside accuracy.

    TR
    1 Source, 15d ago, first seen 15d ago

    TLDR

    The team announced the leaderboard on September 15, 2026, saying it aims to help people choose a model for search-heavy work. It describes tests of 24 frontier models using the same search, with and without web access. The team says its previously internal methodology measures accuracy alongside cost and latency.

    Combined views

    9.2K

    1 Source, first seen 15d ago

    Combined views

    9.2K

    1 Source, first seen 15d ago

    51 likes
    51 likes
    2 comments
    12 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    12 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @travers00today we launched a leaderboard for how frontier models perform on web search tasks "which model should I use for search-heavy work" is a question we get constantly and there weren't good public evals for it. nothing that tests search capability end to end, with cost and latency next to accuracy so we made the methodology we use internally public. 24 models w/ same search, with and without web access

    1 Source

    @travers00today we launched a leaderboard for how frontier models perform on web search tasks "which model should I use for search-heavy work" is a question we get constantly and there weren't good public evals for it. nothing that tests search capability end to end, with cost and latency next to accuracy so we made the methodology we use internally public. 24 models w/ same search, with and without web access