VC-Attention debuts with claimed attention speedups on MiniMax-H3
NunchuxAI says VC-Attention speeds up attention on MiniMax-H3 by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, without retraining.
TLDR
NunchuxAI introduced VC-Attention, a low-bit attention method it says requires no retraining and works with existing sparse attention methods. On MiniMax-H3, it reports attention speedups of 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. The team credits V-Smooth with reducing value quantization error and ExpCast-FP8 with speeding up softmax. NunchuxAI says its proprietary extension, Nunchux Attention, raises those speedups to 1.9× on B200 and 1.8× on B300.
Combined views
59.5K
5 Sources, first seen 1d ago
VC-Attention debuts with claimed attention speedups on MiniMax-H3
NunchuxAI says VC-Attention speeds up attention on MiniMax-H3 by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, without retraining.