Announcement
GLM-5.3 serving reportedly handles 66 sessions per prefill group at 101 tokens per second per user
PrimeIntellect says a 1:4 prefill-to-decode ratio delivered that result on GB200 NVL72 hardware.
TLDR
PrimeIntellect says its GLM-5.3 setup on GB200 NVL72 served 66 sessions per prefill group at 101 tokens per second per user using a 1:4 prefill-to-decode ratio. It also reports roughly five times the prefix-cache capacity with DEP8 versus TEP8 on the same GPUs. NVFP4 compression fit about 50 percent more cached tokens per decoder than FP8, while a vLLM cache layout halved mean KV transfer time.
Combined views
7.1K
5 Sources, first seen ago