Mixed-model AI groups reportedly often fall short of their best single model
A post summarizing an NVIDIA paper says majority voting across copies of the best single model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined.
TLDR
A user’s summary of an NVIDIA paper describes eight model-selection strategies tested across routing, majority voting and LLM-as-judge setups on hard science benchmarks. According to the summary, larger pools of different open models raised theoretical best-case accuracy, but achieved accuracy often fell below the best single model in the pool. Using several copies of one model worked better: majority voting over the best model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined. The summary also says selecting candidates from a single model family gave the largest improvement over a standalone model among the eight strategies.
Combined views
12.9K
2 Sources, first seen 14d ago