Video Delta Net speeds up open-source text-to-video
PhD researcher announces hybrid attention model for real-time text-to-video generation.
TLDR
Haocheng Xi posted about Video Delta Net, a hybrid attention approach for live text-to-video that he says runs faster than playback at near-lossless quality. The post states it accelerates Minimax-H3 and produces 768p video on eight NVIDIA B200 GPUs. Visible replies link to a GitHub repository containing the code and note environment requirements including PyTorch 2.13 plus FlashAttention 4 for the FlexAttention backend. An attached video shows a rapid montage of generated clips. Retweets from engineers and investors highlight the release.
Combined views
209.6K
4 Sources, first seen 28d ago