• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    UMBP reportedly cuts extra KV-cache copies in SGLang

    SemiAnalysis says UMBP's integration in SGLang replaces HiCache's intermediate buffer with a direct path between high-bandwidth memory and DRAM, using a shared pool per node.

    2 Sources, 22d ago, first seen 22d ago

    TLDR

    SemiAnalysis describes HiCache's intermediate host-memory buffer as requiring an extra copy whenever key-value (KV) cache data is loaded or offloaded, with fetches that are not pipelined with computation. It says UMBP provides a direct, layer-wise pipelined path that removes the intermediate copy and stall. The shared pool also deduplicates replicated KV data and allows global placement and eviction decisions across data-parallel ranks and instances, according to SemiAnalysis. The publisher credits AMD's SGLang team, its MoRI (Modular RDMA Interface) library and UMBP's integration for the improvements.

    Combined views

    —

    2 Sources, first seen 22d ago

    Combined views

    —

    2 Sources, first seen 22d ago

    — likes
    — likes
    — comments
    — saves
    — reposts
    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    SemiAnalysis@SemiAnalysis_HiCache sits between HBM and the external DRAM KV store as a per-rank host-side L2 buffer, so every load and offload takes an extra copy, the L3-to-L2 fetch is not pipelined with compute, and under MLA with TP each rank replicates the same KV across PCIe and host memory while making eviction decisions on local information only. UMBP through the SGLang KVCache Store Linker replaces that with a direct, layer-wise pipelined HBM-to-DRAM path into a single shared per-node pool, which removes the intermediate copy and stall, deduplicates the replicated KV, and lets placement and eviction be decided globally across DP ranks and instances. (3/3)22d

    2 Sources

    SemiAnalysis@SemiAnalysis_HiCache sits between HBM and the external DRAM KV store as a per-rank host-side L2 buffer, so every load and offload takes an extra copy, the L3-to-L2 fetch is not pipelined with compute, and under MLA with TP each rank replicates the same KV across PCIe and host memory while making eviction decisions on local information only. UMBP through the SGLang KVCache Store Linker replaces that with a direct, layer-wise pipelined HBM-to-DRAM path into a single shared per-node pool, which removes the intermediate copy and stall, deduplicates the replicated KV, and lets placement and eviction be decided globally across DP ranks and instances. (3/3)22d