• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Posts Highlight Gemini 3.8 Flash Benchmark Scores

    Creators share benchmark tables showing Gemini 3.8 Flash results.

    TK
    AC
    JS
    12 Sources, 28d ago, first seen 28d ago

    TLDR

    Kim Isenberg posted a benchmark table on X claiming Gemini 3.8 Flash outperforms Claude Opus 5, Claude Sonnet 5, GPT-5.6 Sol and GPT-5.6 Terra on Terminal Bench 2.1, HLE and other tests, writing "Google is back." Andrew Curran posted a similar table and stated the model is live for him. Both attached screenshots of the comparisons. Replies on X split between praise for the benchmark gains at lower cost and criticism that real-world performance falls short of frontier models and pricing erases any gains.

    Combined views

    294.6K

    12 Sources, first seen 28d ago

    Combined views

    294.6K

    12 Sources, first seen 28d ago

    3.9K likes
    3.9K likes
    206 comments
    437 saves
    247 reposts
    206 comments
    437 saves
    247 reposts

    Sentiment

    Positive45.4%54.6%Negative

    Summary

    Sentiment

    Positive45.4%54.6%Negative

    Positive accounts welcomed Gemini 3.8 Flash for strong benchmark scores and lower costs, while negative accounts criticized its high hallucination rates and weak real-world reliability.

    Based on 136 sentiment-bearing replies from 130 accounts across 4 conversations.

    Summary

    Positive accounts welcomed Gemini 3.8 Flash for strong benchmark scores and lower costs, while negative accounts criticized its high hallucination rates and weak real-world reliability.

    Based on 136 sentiment-bearing replies from 130 accounts across 4 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    12 Sources

    @kimmonismusGemini 3.8 Flash benchmarks. And holy cow! Flash outperforms 5.6 sol and opus 5 on terminal bench 2.1, HLE and much more. Google is back!
    @AndrewCurran_Gemini 3.8 Flash is live for me now.
    @tkipfit’s a good model (and still insanely fast)
    @rohanpaul_aiGemini 3.8 Flash is out. beats Claude Opus 5 on some really important benchmarks. - 54.9% on HLE-Verified vs 54.4% for Claude Opus 5 - Terminal-Bench 2.1, that test tests whether an agent can actually operate a terminal and successfully finish difficult coding, security, ML, data-science, and systems tasks. 89.4% is essentially a very high task-completion rate under that evaluation setup. - Harvey's Legal Agent Benchmark, Gemini 3.8 Flash scores 10.0% vs Opus 5's 6.7%. The number looks low because this uses an extremely strict all-pass rule: a legal workflow gets credit only when every required criterion passes, including facts, conclusions, citations, structure, and analysis, across complex file-based legal work. Google has not raised the price per token. On harder tasks, Gemini 3.8 Flash may think longer and make more tool calls, so it uses more tokens and the total cost of completing that task can rise even though the token price stays the same.
    @vivnatgemini 3.8 flash is another good model and the science capabilities keep improving :)
    @andrewwhite01RT @vivnat: gemini 3.8 flash is another good model and the science capabilities keep improving :)
    @divy93tSay hello to Gemini 3.8 Flash - it's good at getting stuff done.
    @sunjiao123sun_RT @AndrewCurran_: Gemini 3.8 Flash is live for me now.
    @eigenhectorLmao, I also liked flash 3.8 and have been using it for weeks and was also paid by alphabet 😂
    @ShikharMurtyRT @goodhunt: welp Gemini 3.8 flash is really good

    12 Sources

    @kimmonismusGemini 3.8 Flash benchmarks. And holy cow! Flash outperforms 5.6 sol and opus 5 on terminal bench 2.1, HLE and much more. Google is back!
    @AndrewCurran_Gemini 3.8 Flash is live for me now.
    @tkipfit’s a good model (and still insanely fast)
    @rohanpaul_aiGemini 3.8 Flash is out. beats Claude Opus 5 on some really important benchmarks. - 54.9% on HLE-Verified vs 54.4% for Claude Opus 5 - Terminal-Bench 2.1, that test tests whether an agent can actually operate a terminal and successfully finish difficult coding, security, ML, data-science, and systems tasks. 89.4% is essentially a very high task-completion rate under that evaluation setup. - Harvey's Legal Agent Benchmark, Gemini 3.8 Flash scores 10.0% vs Opus 5's 6.7%. The number looks low because this uses an extremely strict all-pass rule: a legal workflow gets credit only when every required criterion passes, including facts, conclusions, citations, structure, and analysis, across complex file-based legal work. Google has not raised the price per token. On harder tasks, Gemini 3.8 Flash may think longer and make more tool calls, so it uses more tokens and the total cost of completing that task can rise even though the token price stays the same.
    @vivnatgemini 3.8 flash is another good model and the science capabilities keep improving :)
    @andrewwhite01RT @vivnat: gemini 3.8 flash is another good model and the science capabilities keep improving :)
    @divy93tSay hello to Gemini 3.8 Flash - it's good at getting stuff done.
    @sunjiao123sun_RT @AndrewCurran_: Gemini 3.8 Flash is live for me now.
    @eigenhectorLmao, I also liked flash 3.8 and have been using it for weeks and was also paid by alphabet 😂
    @ShikharMurtyRT @goodhunt: welp Gemini 3.8 flash is really good