Uno promises diffusion-style speed with autoregressive LLM quality
Uno’s developer says each layer has separate autoregressive and diffusion weights, enabling parallel sampling without changing the autoregressive output distribution.
TLDR
A talk announcement describes Uno as a drop-in replacement for speculative decoding that needs no separately trained draft model. Its developer claims it is faster than DFlash and EAGLE-3. The announcement scheduled a presentation for September 18, 2026, at 10 a.m. PT / 1 p.m. ET, and the presenter promised new results on faster reinforcement-learning training and comparisons with a retrained DFlash model.
Combined views
9.7K
7 Sources, first seen 2d ago
Uno promises diffusion-style speed with autoregressive LLM quality
Uno’s developer says each layer has separate autoregressive and diffusion weights, enabling parallel sampling without changing the autoregressive output distribution.