• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Pearl-GEMM is claimed to be about 5% faster than Quack on B200 chips

    A former Nvidia researcher says Pearl-GEMM also runs about 42% faster than Flashinfer on B200s, calling Pearl's FP8 mixture-of-experts kernels the fastest for Blackwell.

    sam lessin 🏴‍☠️SL
    Omri WeinsteinOW
    2 Sources, 20d ago, first seen 20d ago

    TLDR

    A former Nvidia researcher claims Pearl has the fastest FP8 kernels for mixture-of-experts AI models on Blackwell chips. The post puts Pearl-GEMM about 5% ahead of Quack and about 42% ahead of Flashinfer on B200s. It describes grouped matrix multiplication as one of AI's most optimized operations and says pushing its performance further is difficult.

    Combined views

    34.8K

    2 Sources, first seen 20d ago

    Combined views

    34.8K

    2 Sources, first seen 20d ago

    234 likes
    234 likes
    15 comments
    57 saves
    53 reposts
    15 comments
    57 saves
    53 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Omri Weinstein@WeinsteinOmriPearl (@prlnet) has the fastest FP8 MoE Kernels for Blackwell chips! Group-Matrix-Multiplication is one of the most optimized operations in AI. As a former researcher at @nvidia I can say firsthand that pushing the performance frontier of MatMuls is... Hard. Pearl-GEMM is ~5% faster than prior state-of-art Quack (@tri_dao) and ~42% faster than Flashinfer on B200s. https://pearlresearch.ai/research/blog/grouped-gemm20d
    sam lessin 🏴‍☠️@lessinRT @WeinsteinOmri: Pearl (@prlnet) has the fastest FP8 MoE Kernels for Blackwell chips! Group-Matrix-Multiplication is one of the most op…20d

    2 Sources

    Omri Weinstein@WeinsteinOmriPearl (@prlnet) has the fastest FP8 MoE Kernels for Blackwell chips! Group-Matrix-Multiplication is one of the most optimized operations in AI. As a former researcher at @nvidia I can say firsthand that pushing the performance frontier of MatMuls is... Hard. Pearl-GEMM is ~5% faster than prior state-of-art Quack (@tri_dao) and ~42% faster than Flashinfer on B200s. https://pearlresearch.ai/research/blog/grouped-gemm20d
    sam lessin 🏴‍☠️@lessinRT @WeinsteinOmri: Pearl (@prlnet) has the fastest FP8 MoE Kernels for Blackwell chips! Group-Matrix-Multiplication is one of the most op…20d