GLM-5.2 Hits 647 Tokens Per Second After Engine Work
Optimizations to TRT-LLM engine produce major inference gains on Grace Blackwell hardware.
TLDR
Yangqing Jia stated the fleet re-engineered the TRT-LLM inference engine end to end for GLM 5.2. The work raised performance from 102 to 647 tokens per second on 2x Grace Blackwell nodes, delivering a 6.3x speedup. Posts from Pete Skomoroch and others framed the result as the new bar for fast inference on the model, with the metric now serving as a performance benchmark for this large language model.
Combined views
8.1K
2 Sources, first seen 66d ago