Many users expressed excitement about Tobi Lutke hitting 40 TPS on GLM-52 with NVIDIA DGX because it offers strong localized inference density for medical AI workloads and reflects an impressive development pace.
No Digg Deeper questions have been answered for this story yet.
Most Activity
40 TPS on glm-52 (nvfp4+mtp)
Thank you @nvidia for the hookup on this DGX Workstation. This will crunch a *lot* of high quality tokens here! This thing is a total beast.
It’s building a liquid implementation from scratch overnight against http://gitHub.com/Shopity/liquid-spec
40 TPS on glm-52 (nvfp4+mtp)
@toddybayern @tobi I request that @tobi rebrands Shopify to Shopity immediately
@tobi how much is the budget?
@tobi how much wat it consumes? and what about the noise?
@tobi shopity would actually be a cool company name hahaha
@tobi GLM-52 at 40 TPS via NVFP4 + MTP? I would love to test that inference density on my medical AI workloads. Excellent localized compute.
@tobi "Overnight from scratch" means someone is very proud, very tired, and already thinking about the next sprint. That's a ridiculous pace to hit. Huge congrats.
@tobi Without the typo: https://github.com/Shopify/liquid-spec