• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Bindu Reddy Says Gemini 3.8 Flash Regresses on Benchmarks

    Bindu Reddy reports Gemini 3.8 Flash underperforms 3.7 Flash on hidden tests.

    RS
    JL
    BR
    6 Sources, 28d ago, first seen 28d ago

    TLDR

    Bindu Reddy, CEO of Abacus.AI, posted that Gemini 3.8 Flash scores lower than 3.7 Flash on the company's benchmark. She wrote that the new version appears overfit to public benchmarks, performs worse on hidden questions, and regresses on data analysis. Reddy called for Google to stop weekly Flash releases and ship Gemini 4.0 instead. In a follow-up she noted frequent checkpoints leave Google trailing open-source efforts. A separate post from Da7_Tech added that 3.7 Flash handles agentic tasks more reliably while 3.8 works better only as a chat model.

    Combined views

    104.7K

    6 Sources, first seen 28d ago

    Combined views

    104.7K

    6 Sources, first seen 28d ago

    881 likes
    881 likes
    143 comments
    119 saves
    47 reposts
    143 comments
    119 saves
    47 reposts

    Sentiment

    Positive27%73%Negative

    Summary

    Sentiment

    Positive27%73%Negative

    Positive accounts praise Gemini 3.8 Flash for strong real-world results and dismiss benchmarks, while negative accounts call it a regression from 3.7 Flash in analysis quality and complex tasks.

    Based on 43 sentiment-bearing replies from 37 accounts across 2 conversations.

    Summary

    Positive accounts praise Gemini 3.8 Flash for strong real-world results and dismiss benchmarks, while negative accounts call it a regression from 3.7 Flash in analysis quality and complex tasks.

    Based on 43 sentiment-bearing replies from 37 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    @bindureddyVery Odd, Gemini keeps releasing new Flash checkpoints every other week - today it's 3.8 Weird to have Google trail behind open-source when Fable and Astra are so close to AGI At the moment, the new Flash 3.8 is not showing much improvement over the last one.
    @Da7_TechI don't get what Google is doing. Gemini 3.7 Flash is way better for agentic tasks than 3.8 Flash. It's calmer, double-checks its work before handing it off, and runs much faster. 3.8 is really only better as a chat model. Why'd they ruin it like this?
    @ScobleizerRT @Da7_Tech: I don't get what Google is doing. Gemini 3.7 Flash is way better for agentic tasks than 3.8 Flash. It's calmer, double-chec…
    @jerryjliu0Gemini 3.8 Flash might be maxxed on other benchmarks, but not quite on document parsing. We benchmarked it on ParseBench. It does slightly better on tables + content faithfulness compared to Gemini 3.7 Flash, but a little worse on charts and semantic formatting. The overall scores are similar to those of 3.7 and 3.6 Flash. We definitely need to evolve the benchmark towards higher difficulty documents. In the meantime though, parsing real-world docs remains a real challenge for frontier models!

    6 Sources

    @bindureddyVery Odd, Gemini keeps releasing new Flash checkpoints every other week - today it's 3.8 Weird to have Google trail behind open-source when Fable and Astra are so close to AGI At the moment, the new Flash 3.8 is not showing much improvement over the last one.
    @Da7_TechI don't get what Google is doing. Gemini 3.7 Flash is way better for agentic tasks than 3.8 Flash. It's calmer, double-checks its work before handing it off, and runs much faster. 3.8 is really only better as a chat model. Why'd they ruin it like this?
    @ScobleizerRT @Da7_Tech: I don't get what Google is doing. Gemini 3.7 Flash is way better for agentic tasks than 3.8 Flash. It's calmer, double-chec…
    @jerryjliu0Gemini 3.8 Flash might be maxxed on other benchmarks, but not quite on document parsing. We benchmarked it on ParseBench. It does slightly better on tables + content faithfulness compared to Gemini 3.7 Flash, but a little worse on charts and semantic formatting. The overall scores are similar to those of 3.7 and 3.6 Flash. We definitely need to evolve the benchmark towards higher difficulty documents. In the meantime though, parsing real-world docs remains a real challenge for frontier models!