The case for text diffusion spreading across language models
An essay argues that refining blocks of tokens could offer advantages for inference, test-time scaling and reinforcement learning.
TLDR
An essay predicts text diffusion will spread across language models as converting existing models gets cheaper. Rather than generating one token at a time, diffusion models refine a noisy block of tokens over several steps. The author argues this could speed up memory-bound generation and offer benefits for test-time computation and reinforcement learning. The DiffusionGemma conversion cited in the essay came with a moderate quality cost.
Combined views
52
1 Source, first seen ago