• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Video Delta Net speeds up open-source text-to-video

    PhD researcher announces hybrid attention model for real-time text-to-video generation.

    HH
    NA
    HX
    4 Sources, 28d ago, first seen 28d ago

    TLDR

    Haocheng Xi posted about Video Delta Net, a hybrid attention approach for live text-to-video that he says runs faster than playback at near-lossless quality. The post states it accelerates Minimax-H3 and produces 768p video on eight NVIDIA B200 GPUs. Visible replies link to a GitHub repository containing the code and note environment requirements including PyTorch 2.13 plus FlashAttention 4 for the FlexAttention backend. An attached video shows a rapid montage of generated clips. Retweets from engineers and investors highlight the release.

    Combined views

    209.6K

    4 Sources, first seen 28d ago

    Combined views

    209.6K

    4 Sources, first seen 28d ago

    1.1K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1.1K likes
    68 comments
    934 saves
    182 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    68 comments
    934 saves
    182 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    @HaochengXiUCBOpen-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality. VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8× NVIDIA B200 GPUs. Checkpoints + training/inference code + Technical Blog ⬇️ (1/6)
    @navalRT @HaochengXiUCB: Open-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN…
    @drisspgBe me: "Ohh wow this looks cool I wonder what their code is like" https://github.com/OpenVDN/vdn-minimax-h3 "Create the environment. We recommend PyTorch 2.13 (torch.__version__ = 2.13.0+cu129) and installing FlashAttention 4, since our code requires FlexAttention's Flash backend." Nice!
    @cHHilleeRT @drisspg: Be me: "Ohh wow this looks cool I wonder what their code is like" https://github.com/OpenVDN/vdn-minimax-h3 "Create the environment. We recomm…

    4 Sources

    @HaochengXiUCBOpen-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality. VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11 seconds on 8× NVIDIA B200 GPUs. Checkpoints + training/inference code + Technical Blog ⬇️ (1/6)
    @navalRT @HaochengXiUCB: Open-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN…
    @drisspgBe me: "Ohh wow this looks cool I wonder what their code is like" https://github.com/OpenVDN/vdn-minimax-h3 "Create the environment. We recommend PyTorch 2.13 (torch.__version__ = 2.13.0+cu129) and installing FlashAttention 4, since our code requires FlexAttention's Flash backend." Nice!
    @cHHilleeRT @drisspg: Be me: "Ohh wow this looks cool I wonder what their code is like" https://github.com/OpenVDN/vdn-minimax-h3 "Create the environment. We recomm…