• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GLM-5.2 Hits 647 Tokens Per Second After Engine Work

    Optimizations to TRT-LLM engine produce major inference gains on Grace Blackwell hardware.

    YJ
    PS
    2 Sources, 66d ago, first seen 66d ago

    TLDR

    Yangqing Jia stated the fleet re-engineered the TRT-LLM inference engine end to end for GLM 5.2. The work raised performance from 102 to 647 tokens per second on 2x Grace Blackwell nodes, delivering a 6.3x speedup. Posts from Pete Skomoroch and others framed the result as the new bar for fast inference on the model, with the metric now serving as a performance benchmark for this large language model.

    Combined views

    8.1K

    2 Sources, first seen 66d ago

    Combined views

    8.1K

    2 Sources, first seen 66d ago

    21 likes
    21 likes
    1 comments
    7 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    7 saves
    1 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    GLM-5.2Yangqing JiaIntent Lab

    2 Sources

    @peteskomorochRT @rachelrapp: The new bar for “fast” on GLM-5.2: 600+ TPS.
    @jiayq1/ The fleet re-engineered the TRT-LLM inference engine and made it the fastest engine we know of for GLM 5.2: 102 to 647 tokens/sec on 2x Grace Blackwell nodes, a 6.3x speedup. This is with the fleet building and implementing optimization end to end.

    2 Sources

    @peteskomorochRT @rachelrapp: The new bar for “fast” on GLM-5.2: 600+ TPS.
    @jiayq1/ The fleet re-engineered the TRT-LLM inference engine and made it the fastest engine we know of for GLM 5.2: 102 to 647 tokens/sec on 2x Grace Blackwell nodes, a 6.3x speedup. This is with the fleet building and implementing optimization end to end.

    Related

    Intent Lab Launches Fleet to Turn Intent Into Production Software

    Yangqing Jia, Caffe creator, unveils team turning natural language intent into verified production systems.