• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    VC-Attention debuts with claimed attention speedups on MiniMax-H3

    NunchuxAI says VC-Attention speeds up attention on MiniMax-H3 by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, without retraining.

    Jun-Yan ZhuJZ
    Song HanSH
    Nunchux AINA
    7 Sources, ,

    TLDR

    NunchuxAI introduced VC-Attention, a low-bit attention method it says requires no retraining and works with existing sparse attention methods. On MiniMax-H3, it reports attention speedups of 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2.

    The team credits V-Smooth with reducing value quantization error and ExpCast-FP8 with speeding up softmax. NunchuxAI says its proprietary extension, Nunchux Attention, raises those speedups to 1.9× on B200 and 1.8× on B300.

    Combined views

    88.3K

    7 Sources, first seen 21d ago

    Combined views

    88.3K

    7 Sources, first seen 21d ago

    550 likes
    21d ago
    first seen 21d ago
    550 likes
    34 comments
    285 saves
    100 reposts
    34 comments
    285 saves
    100 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    7 Sources

    Nunchux AI@NunchuxAIIntroducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.21d
    Muyang Li@lmxyy1999Low-bit attention is promising for fast video generation, but softmax remains a major bottleneck on B200 and B300. Our latest work, VC-Attention, addresses this with ExpCast-FP8: a simple linear mapping to FP8 codes that bypasses the expensive exponential and cast. V-Smooth further improves fidelity by reducing value quantization error. On MiniMax-H3, VC-Attention offers 1.6× attention speedup on B200 and 1.5× on B300 over BF16 FlashAttention4. No retraining needed, and it works with existing sparse attention methods. This work brings us closer to faster, more affordable multimodal inference at @NunchuxAI .21d
    Jun-Yan Zhu@junyanz89RT @NunchuxAI: Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up att…21d
    Song Han@songhan_mitRT @lmxyy1999: Low-bit attention is promising for fast video generation, but softmax remains a major bottleneck on B200 and B300. Our lat…21d
    MiniMax (official)@MiniMax_AIThree paths to faster video attention: compute the same interactions more efficiently, compute fewer in full, or change how information is mixed. Here’s a visual guide. 👇 Thanks to Nunchux AI and collaborators for VC-Attention, bringing training-free low-bit acceleration to MiniMax-H3, with better fidelity than SageAttention2 in the B200 evaluation. The approach balances speed and fidelity: V-Smooth reduces value quantization error, while ExpCast-FP8 makes softmax faster through approximation. Excited to see the community keep building on H3. Could combining low-bit computation with sparse methods like Sol-Attn push efficiency further? We’re looking forward to seeing that explored.20d

    7 Sources

    Nunchux AI@NunchuxAIIntroducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.21d
    Muyang Li@lmxyy1999Low-bit attention is promising for fast video generation, but softmax remains a major bottleneck on B200 and B300. Our latest work, VC-Attention, addresses this with ExpCast-FP8: a simple linear mapping to FP8 codes that bypasses the expensive exponential and cast. V-Smooth further improves fidelity by reducing value quantization error. On MiniMax-H3, VC-Attention offers 1.6× attention speedup on B200 and 1.5× on B300 over BF16 FlashAttention4. No retraining needed, and it works with existing sparse attention methods. This work brings us closer to faster, more affordable multimodal inference at @NunchuxAI .21d
    Jun-Yan Zhu@junyanz89RT @NunchuxAI: Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up att…21d
    Song Han@songhan_mitRT @lmxyy1999: Low-bit attention is promising for fast video generation, but softmax remains a major bottleneck on B200 and B300. Our lat…21d
    MiniMax (official)@MiniMax_AIThree paths to faster video attention: compute the same interactions more efficiently, compute fewer in full, or change how information is mixed. Here’s a visual guide. 👇 Thanks to Nunchux AI and collaborators for VC-Attention, bringing training-free low-bit acceleration to MiniMax-H3, with better fidelity than SageAttention2 in the B200 evaluation. The approach balances speed and fidelity: V-Smooth reduces value quantization error, while ExpCast-FP8 makes softmax faster through approximation. Excited to see the community keep building on H3. Could combining low-bit computation with sparse methods like Sol-Attn push efficiency further? We’re looking forward to seeing that explored.20d