• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Gemini 4 Argon is claimed to guess wrong 15% of the time when it doesn't know an answer

    Merge, citing Artificial Analysis scores from the week of October 5, puts Qwen3.8 Max and Grok 4.7 at 29% on the same measure.

    Zachary NadoZN
    MergeME
    2 Sources, ,

    TLDR

    Merge, citing Artificial Analysis scores for the week of October 5, says Gemini 4 Argon guessed wrong on 15% of AA-Omniscience questions it didn't know, versus 29% for Qwen3.8 Max and Grok 4.7. Merge puts Muse Spark 1.3 at 32% and says GPT-6 Astra had the highest accuracy of the six at 61%, but guessed on 45% of questions it missed. Merge pitches its Gateway as a way to route requests to different models.

    Combined views

    13K

    2 Sources, first seen 13h ago

    Combined views

    13K

    2 Sources, first seen 13h ago

    345 likes
    13h ago
    first seen 13h ago
    345 likes
    5 comments
    34 saves
    29 reposts
    5 comments
    34 saves
    29 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Merge@merge_apiGemini 4 Argon (@GoogleDeepMind) makes up an answer about half as often as the next best frontier model. On AA-Omniscience, when Argon doesn't know something it guesses wrong 15% of the time instead of saying so. Qwen3.8 Max is next at 29%, with Grok 4.7 at 29% and Muse Spark 1.3 at 32%. GPT-6 Astra has the highest accuracy of the six at 61%, but still guesses on 45% of the questions it misses. Whether a wrong answer or no answer costs you more depends on the task. With Merge Gateway you can send each request to the model that fits it. Scores via http://artificialanalysis.ai, week of Oct 5.13h
    Zachary Nado@zacharynadoRT @merge_api: Gemini 4 Argon (@GoogleDeepMind) makes up an answer about half as often as the next best frontier model. On AA-Omniscience,…1h

    2 Sources

    Merge@merge_apiGemini 4 Argon (@GoogleDeepMind) makes up an answer about half as often as the next best frontier model. On AA-Omniscience, when Argon doesn't know something it guesses wrong 15% of the time instead of saying so. Qwen3.8 Max is next at 29%, with Grok 4.7 at 29% and Muse Spark 1.3 at 32%. GPT-6 Astra has the highest accuracy of the six at 61%, but still guesses on 45% of the questions it misses. Whether a wrong answer or no answer costs you more depends on the task. With Merge Gateway you can send each request to the model that fits it. Scores via http://artificialanalysis.ai, week of Oct 5.13h
    Zachary Nado@zacharynadoRT @merge_api: Gemini 4 Argon (@GoogleDeepMind) makes up an answer about half as often as the next best frontier model. On AA-Omniscience,…1h