LLaDA-Image Framework for Strong Image Generators
Paper on arXiv presents a 6B DiT trained from scratch with open recipes.
TLDR
Tanishq Mathew Abraham posted about LLaDA-Image. The linked arXiv paper describes a unified framework that pairs a 6B Diffusion Transformer trained from scratch with a frozen vision-language module built on the LLaDA2.0-Mini backbone. Linked repositories on GitHub and Hugging Face under inclusionAI provide the code and model weights. The abstract states the approach avoids reliance on proprietary components and supplies fully open training recipes.
Combined views
5.4K
1 Source, first seen 27d ago