A blog preview asks how to post-train DiffusionGemma
The post promises a look at what works, what breaks and which supervised fine-tuning (SFT) objective wins.
TLDR
A user describes DiffusionGemma as one of the first large, open-weight uniform diffusion language models and announces a blog about post-training it. The preview promises practical lessons on what works and what breaks, plus what the author calls the winning supervised fine-tuning objective.
Combined views
4
1 Source, first seen 19d ago
reposts