• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Paper Introduces Diffusion-Augmented LLMs

    Tweet shares arXiv paper on diffusion-augmented LLMs for parallel token sampling.

    TM
    BO
    KK
    17 Sources, 27d ago, first seen 27d ago

    TLDR

    Tanishq Mathew Abraham posted about research defining an autoregressive distribution and applying diffusion to sample multiple tokens in parallel. The work presents an 8B Uno model said to outperform the 26B DiffusionGemma. Accompanying materials include the arXiv preprint, a GitHub repository under ifm-ai/uno, a Hugging Face collection named uno, and a project page at s-sahoo.com/uno. The post contains an attached image of the paper.

    Combined views

    76.4K

    17 Sources, first seen 27d ago

    Combined views

    76.4K

    17 Sources, first seen 27d ago

    885 likes
    885 likes
    49 comments
    478 saves
    202 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    49 comments
    478 saves
    202 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    17 Sources

    @ssahoo_Diffusion LLMs have two limitations relative to AR models: (1) Lower quality, and (2) Slower inference at large batch sizes. We address this "Uno" > Retains the AR architecture of LLMs > Each layer has two sets of weights: AR weights and Diffusion weights > Diffusion weights enable parallel sampling from the AR distribution losslessly Results: 💥Faster than all speculative decoding methods: DFlash and EAGLE-3 🔥 Beats ALL diffusion LLMs: Mercury 2, Diffusion Gemma, Llada Links to the paper, models, code below
    @iScienceLuvrUnlocking Lossless Speedups in LLMs via Discrete Diffusion "we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution." "Our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning" website: https://s-sahoo.com/uno/ arxiv: https://arxiv.org/abs/2609.04010 code: https://github.com/ifm-ai/uno huggingface: https://huggingface.co/collections/s-sahoo/uno
    @grahamgcita@ssahoo_ diffusiongemma on you plot should be 3k not 1k.
    @bodonoghue85Seems like interesting work, but "Beats ALL diffusion LLMs" is simply not true. Latency-quality is comfortably Pareto dominated by both Mercury 2 and DiffusionGemma. (I can only post the Uno-Qwen point because strangely the LCBv6 number for Uno is missing from the writeup.....)
    @jwthickstunRT @ssahoo_: Diffusion LLMs have two limitations relative to AR models: (1) Lower quality, and (2) Slower inference at large batch sizes.…
    @dair_aiThis work introduces diffusion-augmented LLMs, a new class of models. They first suggest that speculative decoding needs a separate draft model. Diffusion LLMs give up the quality of the model they replace. This work achieves parallel token generation without either cost. Uno is a class of diffusion augmented LLMs that defines an autoregressive model distribution and uses diffusion to draw multiple tokens in parallel from that same distribution. The parameters split in two. Autoregressive weights train under standard next token prediction, and a lightweight set of diffusion weights is added by a short distillation phase that adds negligible overhead to an existing training pipeline. Because the sampler draws from the AR distribution itself, the speedup is lossless. An existing open-weight AR model can be upgraded rather than retrained. Uno beats leading speculative decoding methods at every evaluated batch size, including the largest the device supports, and reaches up to 3x over the base model. The 8B Uno outperforms the 26B DiffusionGemma and the proprietary Mercury 2 across agentic tool use, coding and long-context reasoning. Paper: https://arxiv.org/abs/2609.04010 Chat with Paper: https://academy.dair.ai/papers/unlocking-lossless-speedups-in-llms-via-discrete-diffusion-2609.04010
    @askalphaxiv“Unlocking Lossless Speedups in LLMs via Discrete Diffusion” Diffusion LMs can generate in parallel but usually sacrifices quality. This paper combines both by training lightweight diffusion adapters to draft multiple tokens in parallel, while the original AR model verifies them exactly. The gives lossless acceleration with no separate draft model, reaching up to 3x faster generation while preserving the base model’s distribution. https://www.alphaxiv.org/abs/2609.04010
    @kastnerkyleRT @askalphaxiv: “Unlocking Lossless Speedups in LLMs via Discrete Diffusion” Diffusion LMs can generate in parallel but usually sacrifice…

    17 Sources

    @ssahoo_Diffusion LLMs have two limitations relative to AR models: (1) Lower quality, and (2) Slower inference at large batch sizes. We address this "Uno" > Retains the AR architecture of LLMs > Each layer has two sets of weights: AR weights and Diffusion weights > Diffusion weights enable parallel sampling from the AR distribution losslessly Results: 💥Faster than all speculative decoding methods: DFlash and EAGLE-3 🔥 Beats ALL diffusion LLMs: Mercury 2, Diffusion Gemma, Llada Links to the paper, models, code below
    @iScienceLuvrUnlocking Lossless Speedups in LLMs via Discrete Diffusion "we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution while using diffusion to draw multiple tokens in parallel from that distribution." "Our 8B Uno model outperforms the leading open d-LLM, the 26B DiffusionGemma, and the proprietary Mercury 2 across all evaluated benchmarks in agentic tool use, coding, and long-context reasoning" website: https://s-sahoo.com/uno/ arxiv: https://arxiv.org/abs/2609.04010 code: https://github.com/ifm-ai/uno huggingface: https://huggingface.co/collections/s-sahoo/uno
    @grahamgcita@ssahoo_ diffusiongemma on you plot should be 3k not 1k.
    @bodonoghue85Seems like interesting work, but "Beats ALL diffusion LLMs" is simply not true. Latency-quality is comfortably Pareto dominated by both Mercury 2 and DiffusionGemma. (I can only post the Uno-Qwen point because strangely the LCBv6 number for Uno is missing from the writeup.....)
    @jwthickstunRT @ssahoo_: Diffusion LLMs have two limitations relative to AR models: (1) Lower quality, and (2) Slower inference at large batch sizes.…
    @dair_aiThis work introduces diffusion-augmented LLMs, a new class of models. They first suggest that speculative decoding needs a separate draft model. Diffusion LLMs give up the quality of the model they replace. This work achieves parallel token generation without either cost. Uno is a class of diffusion augmented LLMs that defines an autoregressive model distribution and uses diffusion to draw multiple tokens in parallel from that same distribution. The parameters split in two. Autoregressive weights train under standard next token prediction, and a lightweight set of diffusion weights is added by a short distillation phase that adds negligible overhead to an existing training pipeline. Because the sampler draws from the AR distribution itself, the speedup is lossless. An existing open-weight AR model can be upgraded rather than retrained. Uno beats leading speculative decoding methods at every evaluated batch size, including the largest the device supports, and reaches up to 3x over the base model. The 8B Uno outperforms the 26B DiffusionGemma and the proprietary Mercury 2 across agentic tool use, coding and long-context reasoning. Paper: https://arxiv.org/abs/2609.04010 Chat with Paper: https://academy.dair.ai/papers/unlocking-lossless-speedups-in-llms-via-discrete-diffusion-2609.04010
    @askalphaxiv“Unlocking Lossless Speedups in LLMs via Discrete Diffusion” Diffusion LMs can generate in parallel but usually sacrifices quality. This paper combines both by training lightweight diffusion adapters to draft multiple tokens in parallel, while the original AR model verifies them exactly. The gives lossless acceleration with no separate draft model, reaching up to 3x faster generation while preserving the base model’s distribution. https://www.alphaxiv.org/abs/2609.04010
    @kastnerkyleRT @askalphaxiv: “Unlocking Lossless Speedups in LLMs via Discrete Diffusion” Diffusion LMs can generate in parallel but usually sacrifice…