• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    DiffusionGemma reportedly gets a 3–10x speedup by fixing structured response tokens in place

    A developer says their diffgemma implementation reduces the work needed to generate answers and suggests the technique could probably also be used in vLLM.

    MM
    1 Source, 14d ago, first seen 14d ago

    TLDR

    A developer reports speeding up Google's DiffusionGemma 3–10x by fixing structured response tokens in place. They say the model already works with probabilities, and holding those tokens in place reduces the work needed to produce answers. The developer says they implemented the approach in diffgemma and suggests it could probably make its way into vLLM as well.

    Combined views

    11.5K

    1 Source, first seen 14d ago

    Combined views

    11.5K

    1 Source, first seen 14d ago

    216 likes
    216 likes
    11 comments
    194 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    194 saves
    17 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @mmastracIt turns out that you can speed up Google's @googlegemma 's DiffusionGemma 3-10x by borrowing some of Jev's ideas. The @GoogleDeepMind model is working off probabilities already - if you fix the structured tokens of the response in place, you can drastically reduce the amount of work you do and get answers in far less time. I implemented this for diffgemma, but it could probably find its way into vLLM as well: https://github.com/mmastrac/diffgemma/pull/21

    1 Source

    @mmastracIt turns out that you can speed up Google's @googlegemma 's DiffusionGemma 3-10x by borrowing some of Jev's ideas. The @GoogleDeepMind model is working off probabilities already - if you fix the structured tokens of the response in place, you can drastically reduce the amount of work you do and get answers in far less time. I implemented this for diffgemma, but it could probably find its way into vLLM as well: https://github.com/mmastrac/diffgemma/pull/21