Nvidia's TwELL approach reportedly speeds up sparse LLMs without accuracy loss
The post claims over 20% speedups in inference and training on H100 GPUs, plus lower peak memory consumption and energy use.
TLDR
A user describes a paper by Nvidia and researchers introducing TwELL, a custom sparse packing format paired with specialized GPU code. The approach tackles the overhead of skipping inactive model neurons: according to the post, it routes 99% of sparse tokens through a fast path, with a small dense backup matrix for heavy outliers. The user claims test models achieved over 99% sparsity without accuracy loss, alongside over 20% speedups in forward inference and training on H100 GPUs.
Combined views
15
1 Source, first seen 16d ago