• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    DeepSeek V4.1-Flash is claimed to cut global KV to 890 bytes per token and shrink its persistent cache by 8x

    A post says the model stays competitive on agentic benchmarks and names three techniques behind the claimed cache reductions.

    KD
    1 Source, 5h ago, first seen 5h ago

    TLDR

    A user claims DeepSeek V4.1-Flash cuts global KV to 890 bytes per token and shrinks its persistent cache by 8x while staying competitive on agentic benchmarks. The post names causal encoder-decoder, compressed sparse attention and SWA bounded replay as key techniques, and links a video.

    Combined views

    23

    1 Source, first seen 5h ago

    reposts

    Combined views

    23

    1 Source, first seen 5h ago

    4 reposts
    4

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @CSProfKGDRT @jbhuang0604: DeepSeek V4.1-Flash cuts the global KV to 890 bytes per token and shrinks the persistent cache by 8x, all while staying co…

    1 Source

    @CSProfKGDRT @jbhuang0604: DeepSeek V4.1-Flash cuts the global KV to 890 bytes per token and shrinks the persistent cache by 8x, all while staying co…