Report
Gemini 4 Argon is claimed to guess wrong 15% of the time when it doesn't know an answer
Merge, citing Artificial Analysis scores from the week of October 5, puts Qwen3.8 Max and Grok 4.7 at 29% on the same measure.
TLDR
Merge, citing Artificial Analysis scores for the week of October 5, says Gemini 4 Argon guessed wrong on 15% of AA-Omniscience questions it didn't know, versus 29% for Qwen3.8 Max and Grok 4.7. Merge puts Muse Spark 1.3 at 32% and says GPT-6 Astra had the highest accuracy of the six at 61%, but guessed on 45% of questions it missed. Merge pitches its Gateway as a way to route requests to different models.
Combined views
13K
2 Sources, first seen ago
