• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Nvidia Executive Claims Record Groq Inference Speed

    Executive at Hot Chips cites slides showing Groq 3 LPX hardware speed.

    B(
    SM
    TK
    4 Sources, 37d ago, first seen 37d ago

    TLDR

    A post by @firstadopter reports an Nvidia executive at Hot Chips presented slides claiming Groq 3 LPX achieved 10,996 tokens per second on Gemma 4 31B. Another post references an Artificial Analysis benchmark of the same hardware recording 3,431 tokens per second output at 100k context length, with similar rates across 10k and 100k input tests. Both posts treat the figures as record results for AI inference.

    Combined views

    47K

    4 Sources, first seen 37d ago

    Combined views

    47K

    4 Sources, first seen 37d ago

    375 likes
    375 likes
    17 comments
    43 saves
    15 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    17 comments
    43 saves
    15 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    @sundeep⚡️⚡️⚡️
    @firstadopter"fastest AI inference on the planet" - Nvidia executive at Hot Chips
    @beffjezosJust in at Hot Chips: CUDA coming to Groq LPUs
    @EpochAIResearchNvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s), according to a @ArtificialAnlys benchmarks of Gemma 4 31B. LPX does this by using a small amount (128 GB) of ultrafast SRAM in place of HBM. Nvidia proposes combining LPUs with GPUs to bring this speed to frontier-scale models. A look at how that (might) work below.

    4 Sources

    @sundeep⚡️⚡️⚡️
    @firstadopter"fastest AI inference on the planet" - Nvidia executive at Hot Chips
    @beffjezosJust in at Hot Chips: CUDA coming to Groq LPUs
    @EpochAIResearchNvidia's Groq 3 LPX system excels at fast decode of small models (>3,400 tok/s), according to a @ArtificialAnlys benchmarks of Gemma 4 31B. LPX does this by using a small amount (128 GB) of ultrafast SRAM in place of HBM. Nvidia proposes combining LPUs with GPUs to bring this speed to frontier-scale models. A look at how that (might) work below.