• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Mixed-model AI groups reportedly often fall short of their best single model

    A post summarizing an NVIDIA paper says majority voting across copies of the best single model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined.

    EL
    2 Sources, 14d ago, first seen 14d ago

    TLDR

    A user’s summary of an NVIDIA paper describes eight model-selection strategies tested across routing, majority voting and LLM-as-judge setups on hard science benchmarks. According to the summary, larger pools of different open models raised theoretical best-case accuracy, but achieved accuracy often fell below the best single model in the pool. Using several copies of one model worked better: majority voting over the best model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined. The summary also says selecting candidates from a single model family gave the largest improvement over a standalone model among the eight strategies.

    Combined views

    12.9K

    2 Sources, first seen 14d ago

    Combined views

    12.9K

    2 Sources, first seen 14d ago

    160 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    160 likes
    19 comments
    166 saves
    28 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    19 comments
    166 saves
    28 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @omarsar0Banger paper from NVIDIA. It's on the topic of choosing which models go into a multi-agent system. The team compared eight selection strategies, based on size, accuracy, answer diversity and error diversity, across routing, majority vote and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy. Achieved accuracy often fell below the single best model in the pool. Using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined. Choosing candidates from a single model family gave the largest improvement over a standalone model of all eight strategies. Before adding another model to a router or ensemble, measure what it adds. Paper: https://arxiv.org/abs/2609.17306 Chat with Paper: https://academy.dair.ai/papers/mo-models-mo-problems-how-to-best-select-model-pools-when-designing-multi-agent-2609.17306

    2 Sources

    @omarsar0Banger paper from NVIDIA. It's on the topic of choosing which models go into a multi-agent system. The team compared eight selection strategies, based on size, accuracy, answer diversity and error diversity, across routing, majority vote and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy. Achieved accuracy often fell below the single best model in the pool. Using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined. Choosing candidates from a single model family gave the largest improvement over a standalone model of all eight strategies. Before adding another model to a router or ensemble, measure what it adds. Paper: https://arxiv.org/abs/2609.17306 Chat with Paper: https://academy.dair.ai/papers/mo-models-mo-problems-how-to-best-select-model-pools-when-designing-multi-agent-2609.17306