DiffusionGemma blog announcement promises a look at what works in post-training
The author describes DiffusionGemma as one of the first large, open-weight uniform diffusion language models and asks how the community can post-train it.
TLDR
The announcement says a new blog explores post-training DiffusionGemma—what works, what breaks and which supervised fine-tuning (SFT) objective wins.
Combined views
9.6K
1 Source, first seen 19d ago
101 likes