Turbo-dLLM debuts as an open-source library for training diffusion language models
The Turbo-dLLM team reports DFlash2 speculative-decoder training speedups of 2.48x at 512K context length and 7.59x at 1M context length on eight H100 GPUs.
TLDR
Turbo-dLLMβs team announced an open-source library for training diffusion language models at scale. The library introduces Context-Sharded Block Parallelism, a strategy the team says improves training efficiency, especially as context length increases. For DFlash2 speculative-decoder training on eight H100 GPUs, the team reports speedups of 2.48x at 512K context length and 7.59x at 1M context length.
Combined views
11.4K
4 Sources, first seen ago