• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    GLM-5.3 serving reportedly handles 66 sessions per prefill group at 101 tokens per second per user

    PrimeIntellect says a 1:4 prefill-to-decode ratio delivered that result on GB200 NVL72 hardware.

    PI
    5 Sources, 2h ago, first seen 2h ago

    TLDR

    PrimeIntellect says its GLM-5.3 setup on GB200 NVL72 served 66 sessions per prefill group at 101 tokens per second per user using a 1:4 prefill-to-decode ratio. It also reports roughly five times the prefix-cache capacity with DEP8 versus TEP8 on the same GPUs. NVFP4 compression fit about 50 percent more cached tokens per decoder than FP8, while a vLLM cache layout halved mean KV transfer time.

    Combined views

    7.1K

    5 Sources, first seen 2h ago

    Combined views

    7.1K

    5 Sources, first seen 2h ago

    196 likes
    196 likes
    6 comments
    14 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    6 comments
    14 saves
    2 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @PrimeIntellectLong-context agent serving depends on retaining history, scheduling new work, and moving cached state efficiently. We optimized these paths separately: 1. Prefill topology and scheduling 2. Compressed KV and a fused attention kernel 3. Transfer-friendly cache layout2h

    5 Sources

    @PrimeIntellectLong-context agent serving depends on retaining history, scheduling new work, and moving cached state efficiently. We optimized these paths separately: 1. Prefill topology and scheduling 2. Compressed KV and a fused attention kernel 3. Transfer-friendly cache layout2h