Gemini 3.8 Flash and Muse Spark 1.3 are “benchmaxxed,” SemiAnalysis says
SemiAnalysis says the two models are comparable to GPT-6 and Fable 5.1 on Terminal Bench 2.1, but their Terminal Bench 4.0 performance is markedly worse.
TLDR
SemiAnalysis calls Gemini 3.8 Flash and Muse Spark 1.3 two of the most clearly “benchmaxxed” models it has seen—a criticism of excessive optimization for benchmark scores. It points to a contrast: both are comparable to GPT-6 and Fable 5.1 on Terminal Bench 2.1, but show markedly worse performance on Terminal Bench 4.0.
Combined views
647.6K
8 Sources, first seen 23d ago