Nvidia Executive Claims Record Groq Inference Speed
Executive at Hot Chips cites slides showing Groq 3 LPX hardware speed.
A post by @firstadopter reports an Nvidia executive at Hot Chips presented slides claiming Groq 3 LPX achieved 10,996 tokens per second on Gemma 4 31B. Another post references an Artificial Analysis benchmark of the same hardware recording 3,431 tokens per second output at 100k context length, with similar rates across 10k and 100k input tests. Both posts treat the figures as record results for AI inference.
Combined views
13.5K
3 posts, first seen 1d ago


