Agentic inference is making KV cache management and data movement increasingly central to the serving stack. The new AgentX / InferenceXv3 analysis from @SemiAnalysis_ captures this shift well. Glad to see Mooncake highlighted for external KV caching and disaggregated P/D data…
AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing? $3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355, B200