• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Nvidia research is claimed to make language models faster without losing accuracy

    A user says a sparse data format called TwELL, paired with custom GPU code, delivered over 20% speedups in both inference and training on Nvidia H100s.

    SU
    1 Source, 18d ago, first seen 18d ago

    TLDR

    A user describes research by Nvidia and collaborators targeting a problem with sparse language models: skipping inactive neurons can create irregular memory access that slows GPUs down. The post says TwELL and custom CUDA kernels address this with a fast processing path and a dense backup matrix for heavy outliers. It claims test models achieved over 99% sparsity without losing accuracy, alongside over 20% speedups in inference and training on H100 GPUs and reductions in peak memory consumption and energy use.

    Combined views

    8.9K

    1 Source, first seen 18d ago

    Combined views

    8.9K

    1 Source, first seen 18d ago

    249 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    249 likes
    5 comments
    214 saves
    39 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 comments
    214 saves
    39 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @thesupermannxNvidia proved you can make transformer LLMs sparser, faster, and lighter at the same time, without losing accuracy. Human brain is wildly efficient because it only fires the specific neurons it needs for a thought. Large Language Models naturally try to do the exact same thing, over 95% of the neurons in an LLM's feed-forward layers stay completely silent for any given word. They do almost no math on a day-to-day basis. So why don't they run instantly? Because modern GPUs are built for brute-force, predictable blocks of dense math. When a model tries to skip inactive neurons, it creates random, unstructured gaps. GPUs hate irregular memory access. Historically, trying to force a GPU to run a sparse model actually made it slower. The overhead of managing the empty space ate all the savings. Now, Nvidia and researchers solved it. They dropped a new paper introducing a custom sparse packing format called TwELL paired with brand-new CUDA kernels built directly for modern hardware. Instead of forcing the GPU to choke on irregular gaps, they built a hybrid system: routing 99% of sparse tokens through a hyper-fast path, while using a tiny dense backup matrix as a safety valve for heavy outliers. The results are staggering: • Simple L1 regularization forced over 99% sparsity in test models with zero loss in downstream intelligence. • Over 20% direct speedups in both forward inference and training on H100s. • Massive drops in peak memory consumption and energy usage. • Fully open-source code and kernels ready to deploy. For years, the industry thought the only way to make models faster was to make them smaller or quantize them into lower bits. Nvidia unlocked a completely different dimension.

    1 Source

    @thesupermannxNvidia proved you can make transformer LLMs sparser, faster, and lighter at the same time, without losing accuracy. Human brain is wildly efficient because it only fires the specific neurons it needs for a thought. Large Language Models naturally try to do the exact same thing, over 95% of the neurons in an LLM's feed-forward layers stay completely silent for any given word. They do almost no math on a day-to-day basis. So why don't they run instantly? Because modern GPUs are built for brute-force, predictable blocks of dense math. When a model tries to skip inactive neurons, it creates random, unstructured gaps. GPUs hate irregular memory access. Historically, trying to force a GPU to run a sparse model actually made it slower. The overhead of managing the empty space ate all the savings. Now, Nvidia and researchers solved it. They dropped a new paper introducing a custom sparse packing format called TwELL paired with brand-new CUDA kernels built directly for modern hardware. Instead of forcing the GPU to choke on irregular gaps, they built a hybrid system: routing 99% of sparse tokens through a hyper-fast path, while using a tiny dense backup matrix as a safety valve for heavy outliers. The results are staggering: • Simple L1 regularization forced over 99% sparsity in test models with zero loss in downstream intelligence. • Over 20% direct speedups in both forward inference and training on H100s. • Massive drops in peak memory consumption and energy usage. • Fully open-source code and kernels ready to deploy. For years, the industry thought the only way to make models faster was to make them smaller or quantize them into lower bits. Nvidia unlocked a completely different dimension.