Fixing structured tokens reportedly speeds up DiffusionGemma 3–10×
A developer says they implemented the technique in diffgemma, holding a response’s structured tokens in place to reduce the work needed to generate answers.
TLDR
A developer reports a 3–10× speedup for Google’s DiffusionGemma by fixing a response’s structured tokens in place. They say the model already works with probabilities, and fixing those tokens reduces the work it needs to do. The developer says they implemented the approach in diffgemma and suggests it could probably make its way into vLLM as well.
Combined views
2
1 Source, first seen 14d ago