• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    LLaDA-Image Framework for Strong Image Generators

    Paper on arXiv presents a 6B DiT trained from scratch with open recipes.

    TM
    1 Source, 27d ago, first seen 27d ago

    TLDR

    Tanishq Mathew Abraham posted about LLaDA-Image. The linked arXiv paper describes a unified framework that pairs a 6B Diffusion Transformer trained from scratch with a frozen vision-language module built on the LLaDA2.0-Mini backbone. Linked repositories on GitHub and Hugging Face under inclusionAI provide the code and model weights. The abstract states the approach avoids reliance on proprietary components and supplies fully open training recipes.

    Combined views

    5.4K

    1 Source, first seen 27d ago

    Combined views

    5.4K

    1 Source, first seen 27d ago

    82 likes
    82 likes
    3 comments
    67 saves
    13 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    67 saves
    13 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @iScienceLuvrLLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes "We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98% of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes." abs: https://arxiv.org/abs/2609.03796 code: https://github.com/inclusionAI/LLaDA-Image huggingface: https://huggingface.co/inclusionAI/LLaDA-Image

    1 Source

    @iScienceLuvrLLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes "We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98% of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes." abs: https://arxiv.org/abs/2609.03796 code: https://github.com/inclusionAI/LLaDA-Image huggingface: https://huggingface.co/inclusionAI/LLaDA-Image