Nvidia research is claimed to make language models faster without losing accuracy
A user says a sparse data format called TwELL, paired with custom GPU code, delivered over 20% speedups in both inference and training on Nvidia H100s.
TLDR
A user describes research by Nvidia and collaborators targeting a problem with sparse language models: skipping inactive neurons can create irregular memory access that slows GPUs down. The post says TwELL and custom CUDA kernels address this with a fast processing path and a dense backup matrix for heavy outliers. It claims test models achieved over 99% sparsity without losing accuracy, alongside over 20% speedups in inference and training on H100 GPUs and reductions in peak memory consumption and energy use.
Combined views
8.9K
1 Source, first seen 18d ago