DiffusionGemma reportedly gets a 3–10x speedup by fixing structured response tokens in place
A developer says their diffgemma implementation reduces the work needed to generate answers and suggests the technique could probably also be used in vLLM.
TLDR
A developer reports speeding up Google's DiffusionGemma 3–10x by fixing structured response tokens in place. They say the model already works with probabilities, and holding those tokens in place reduces the work needed to produce answers. The developer says they implemented the approach in diffgemma and suggests it could probably make its way into vLLM as well.
Combined views
11.5K
1 Source, first seen 14d ago