• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Braintrust says web search narrowed the gap between AI models on current-events questions

    Braintrust evaluated 1,329 questions across four models and 14 conditions. It reports bigger search gains for more recent events, while a fifth search query was associated with lower performance.

    BR
    1 Source, 30d ago, first seen 30d ago

    TLDR

    Braintrust says it compared You.com, providers’ built-in search and no search on 1,329 current-events questions across four models and 14 conditions. Search reduced the gap between models from 47.9 points to 5.6, it reports. Gains from search declined with event age, from about 45 points for recent events to about 24 for the oldest. Braintrust also reports that runs with five or more searches scored 19–49%, with a fifth query associated with lower performance.

    Combined views

    17.6K

    1 Source, first seen 30d ago

    Combined views

    17.6K

    1 Source, first seen 30d ago

    45 likes
    45 likes
    5 comments
    29 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 comments
    29 saves
    10 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @braintrustImagine doing your job without ever looking anything up on the internet. That's an agent without web search. Giving agents the web changed what they can do, especially on anything recent that isn't reflected in training data. But agents don't search like people, so optimizing their performance is a new challenge. There's a lot of great research out there on search behavior, but we wanted to answer a more operational question: when should search be on, and how should you configure it? We evaluated 1,329 current events questions across 4 models and 14 conditions, comparing @youdotcom, provider built-in search, and no search. We found that: - Search reduced the gap between models from 47.9 points to 5.6 - Retrieval gain declined with event age, from ~45 points for recent events to ~24 for the oldest - Runs with 5+ searches scored 19–49%. A fifth query was associated with lower performance Read the research → https://braintrustdata.link/web-search-eval

    1 Source

    @braintrustImagine doing your job without ever looking anything up on the internet. That's an agent without web search. Giving agents the web changed what they can do, especially on anything recent that isn't reflected in training data. But agents don't search like people, so optimizing their performance is a new challenge. There's a lot of great research out there on search behavior, but we wanted to answer a more operational question: when should search be on, and how should you configure it? We evaluated 1,329 current events questions across 4 models and 14 conditions, comparing @youdotcom, provider built-in search, and no search. We found that: - Search reduced the gap between models from 47.9 points to 5.6 - Retrieval gain declined with event age, from ~45 points for recent events to ~24 for the oldest - Runs with 5+ searches scored 19–49%. A fifth query was associated with lower performance Read the research → https://braintrustdata.link/web-search-eval