Turbo-dLLM debuts as an open-source library for training diffusion language models
The Turbo-dLLM team reports DFlash2 speculative-decoder training speedups of 2.48x at 512K context length and 7.59x at 1M context length on eight H100 GPUs.
TLDR
Turbo-dLLM’s team announced an open-source library for training diffusion language models at scale. The library introduces Context-Sharded Block Parallelism, a strategy the team says improves training efficiency, especially as context length increases. For DFlash2 speculative-decoder training on eight H100 GPUs, the team reports speedups of 2.48x at 512K context length and 7.59x at 1M context length.
Combined views
54.9K
3 Sources, first seen 17h ago
Turbo-dLLM debuts as an open-source library for training diffusion language models
The Turbo-dLLM team reports DFlash2 speculative-decoder training speedups of 2.48x at 512K context length and 7.59x at 1M context length on eight H100 GPUs.