• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Gemini 3.8 Flash and Muse Spark 1.3 are “benchmaxxed,” SemiAnalysis says

    SemiAnalysis says the two models are comparable to GPT-6 and Fable 5.1 on Terminal Bench 2.1, but their Terminal Bench 4.0 performance is markedly worse.

    AW
    T(
    NR
    8 Sources, ,

    TLDR

    SemiAnalysis calls Gemini 3.8 Flash and Muse Spark 1.3 two of the most clearly “benchmaxxed” models it has seen—a criticism of excessive optimization for benchmark scores. It points to a contrast: both are comparable to GPT-6 and Fable 5.1 on Terminal Bench 2.1, but show markedly worse performance on Terminal Bench 4.0.

    Combined views

    647.6K

    8 Sources, first seen 23d ago

    Combined views

    647.6K

    8 Sources, first seen 23d ago

    5.1K likes
    23d ago
    first seen 23d ago
    5.1K likes
    197 comments
    755 saves
    211 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    197 comments
    755 saves
    211 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    8 Sources

    @SemiAnalysis_Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵
    @alexandr_wangthis is a silly argument. GPT-5.6 Sol was 88.8% on TerminalBench 2.1 and 37.3% on TerminalBench 4.0. That doesn't make it a more benchmaxxed model than GPT-6 Astra even though the diff is larger. TerminalBench 4.0 is simply a much harder eval that isn't saturated, whereas TerminalBench 2.1 is. We don't claim Muse Spark 1.3 is as strong as Astra or Fable 5.1, but it is significantly more cost-effective. Our future models will compete more directly with those models.
    @firstadopterGemini's reputation for benchmaxxing is unmatched. Where is 3.5 Pro?
    @teortaxesTexBut how does "saturation" happen, Wang? For the record, GLM 5.3 Flash destroys Muse Spark (and much else) on TB 4.0. I am agnostic as to whether it's genuinely better, as I haven't used Spark.
    @nazneenrajani@SemiAnalysis_ We have 3 held out benchmarks where even the leading model is <50% Cybersecurity benchmark https://CWE-bench.com CUA world Science Benchmark

    8 Sources

    @SemiAnalysis_Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵
    @alexandr_wangthis is a silly argument. GPT-5.6 Sol was 88.8% on TerminalBench 2.1 and 37.3% on TerminalBench 4.0. That doesn't make it a more benchmaxxed model than GPT-6 Astra even though the diff is larger. TerminalBench 4.0 is simply a much harder eval that isn't saturated, whereas TerminalBench 2.1 is. We don't claim Muse Spark 1.3 is as strong as Astra or Fable 5.1, but it is significantly more cost-effective. Our future models will compete more directly with those models.
    @firstadopterGemini's reputation for benchmaxxing is unmatched. Where is 3.5 Pro?
    @teortaxesTexBut how does "saturation" happen, Wang? For the record, GLM 5.3 Flash destroys Muse Spark (and much else) on TB 4.0. I am agnostic as to whether it's genuinely better, as I haven't used Spark.
    @nazneenrajani@SemiAnalysis_ We have 3 held out benchmarks where even the leading model is <50% Cybersecurity benchmark https://CWE-bench.com CUA world Science Benchmark