Nvidia Rubin vLLM inference claimed to have 3.2x better profit per gigawatt than GB300 NVL72
SemiAnalysis also claims up to 10x better performance per dollar than GB300 NVL72 on the vLLM engine.
TLDR
SemiAnalysis claims Nvidia Rubin inference running on vLLM offers 3.2x better profit per gigawatt and up to 10x better performance per dollar than GB300 NVL72. It describes vLLM as a widely used production LLM engine.
Combined views
—
1 Source, first seen ago
— likes— comments— saves— reposts
