• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Uno promises diffusion-style speed with autoregressive LLM quality

    Uno’s developer says each layer has separate autoregressive and diffusion weights, enabling parallel sampling without changing the autoregressive output distribution.

    EmadEM
    Yuntian DengYD
    Leo BoytsovLB
    9 Sources, ,

    TLDR

    A talk announcement describes Uno as a drop-in replacement for speculative decoding that needs no separately trained draft model. Its developer claims it is faster than DFlash and EAGLE-3. The announcement scheduled a presentation for September 18, 2026, at 10 a.m. PT / 1 p.m. ET, and the presenter promised new results on faster reinforcement-learning training and comparisons with a retrained DFlash model.

    Combined views

    230K

    9 Sources, first seen 22d ago

    Combined views

    230K

    9 Sources, first seen 22d ago

    1.1K likes
    22d ago
    first seen 22d ago
    1.1K likes
    82 comments
    699 saves
    383 reposts
    82 comments
    699 saves
    383 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    9 Sources

    Discrete Diffusion Reading Group@diffusion_llmsThis week, @ssahoo_ presents diffusion-augmented LLMs: 🔥 AR-level Quality ⚡️ Diffusion-level speed ✨ A drop-in replacement for speculative decoding: >> No separate draft model to train, plus >> Faster RL post-training. The method, Uno, beats all diffusion models by a wide margin in system throughput, agentic evals, and long-context reasoning. 📅 Sep 18, Friday 🕒 10am PT / 1PM ET22d
    Subham Sahoo@ssahoo_I’ll share new experimental results in this talk, including: ⚡️ Faster RL training 📊 Comparisons with a retrained DFlash model 📅 Sep 18, Friday (10am PT)22d
    Yuntian Deng@yuntiandengRT @ssahoo_: I’ll share new experimental results in this talk, including: ⚡️ Faster RL training 📊 Comparisons with a retrained DFlash model…22d
    John Thickstun@jwthickstunRT @diffusion_llms: This week, @ssahoo_ presents diffusion-augmented LLMs: 🔥 AR-level Quality ⚡️ Diffusion-level speed ✨ A drop-in replace…21d
    Institute of Foundation Models@IFM_AIToday’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck. We're introducing Uno, a diffusion-augmented LLM that delivers autoregressive quality at diffusion speed. It’s a lossless speedup method that accelerates generation without degrading response quality. With Uno, K2-Horizon-7B outperforms state-of-the-art diffusion methods in both quality and throughput, delivering up to a 2.2× speedup with no loss in quality. Paper: https://arxiv.org/abs/2609.04010 Model available at: https://huggingface.co/IFM/K2-Horizon-7B-Uno20d
    Eric Xing@ericxingDiffusion can multiply LLM, not just add to it to work in parallel, the outcome is a lossless acceleration of any existing LLMs. Try it out!20d
    Hector Liu@waterluffyThis is one of our initial attempts of using applying diffusion models. The plug-and-play property of Uno make it flexible in many areas, think about using them in post training or other complex scenarios. Further, think about it this way: diffusion enhances both depths (denoising steps) and length (parallel testing scaling), it may has the potential to unlock stronger reasoning.20d
    Leo Boytsov@srchvrsRT @waterluffy: This is one of our initial attempts of using applying diffusion models. The plug-and-play property of Uno make it flexible…20d
    Emad@EMostaqueRT @IFM_AI: Today’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck. W…19d

    9 Sources

    Discrete Diffusion Reading Group@diffusion_llmsThis week, @ssahoo_ presents diffusion-augmented LLMs: 🔥 AR-level Quality ⚡️ Diffusion-level speed ✨ A drop-in replacement for speculative decoding: >> No separate draft model to train, plus >> Faster RL post-training. The method, Uno, beats all diffusion models by a wide margin in system throughput, agentic evals, and long-context reasoning. 📅 Sep 18, Friday 🕒 10am PT / 1PM ET22d
    Subham Sahoo@ssahoo_I’ll share new experimental results in this talk, including: ⚡️ Faster RL training 📊 Comparisons with a retrained DFlash model 📅 Sep 18, Friday (10am PT)22d
    Yuntian Deng@yuntiandengRT @ssahoo_: I’ll share new experimental results in this talk, including: ⚡️ Faster RL training 📊 Comparisons with a retrained DFlash model…22d
    John Thickstun@jwthickstunRT @diffusion_llms: This week, @ssahoo_ presents diffusion-augmented LLMs: 🔥 AR-level Quality ⚡️ Diffusion-level speed ✨ A drop-in replace…21d
    Institute of Foundation Models@IFM_AIToday’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck. We're introducing Uno, a diffusion-augmented LLM that delivers autoregressive quality at diffusion speed. It’s a lossless speedup method that accelerates generation without degrading response quality. With Uno, K2-Horizon-7B outperforms state-of-the-art diffusion methods in both quality and throughput, delivering up to a 2.2× speedup with no loss in quality. Paper: https://arxiv.org/abs/2609.04010 Model available at: https://huggingface.co/IFM/K2-Horizon-7B-Uno20d
    Eric Xing@ericxingDiffusion can multiply LLM, not just add to it to work in parallel, the outcome is a lossless acceleration of any existing LLMs. Try it out!20d
    Hector Liu@waterluffyThis is one of our initial attempts of using applying diffusion models. The plug-and-play property of Uno make it flexible in many areas, think about using them in post training or other complex scenarios. Further, think about it this way: diffusion enhances both depths (denoising steps) and length (parallel testing scaling), it may has the potential to unlock stronger reasoning.20d
    Leo Boytsov@srchvrsRT @waterluffy: This is one of our initial attempts of using applying diffusion models. The plug-and-play property of Uno make it flexible…20d
    Emad@EMostaqueRT @IFM_AI: Today’s LLMs still write like typewriters: one token at a time. This sequential process creates a hard inference bottleneck. W…19d