• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    KV cache’s role in AI agent costs

    The comment points to DeepSeek’s compression to roughly 890 bytes per token, arguing that it lets agents keep more context ready for reuse.

    AV
    3 Sources, 18d ago, first seen 18d ago

    TLDR

    A post relays a comment calling KV cache the “quiet cost secret” for AI agents. The comment cites DeepSeek’s compression to roughly 890 bytes per token as a way to keep more context ready for reuse, and claims cache hits versus misses account for most of the bill.

    Combined views

    1.9K

    3 Sources, first seen 18d ago

    likes

    Combined views

    1.9K

    3 Sources, first seen 18d ago

    10 likes
    10
    5 comments
    1 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 comments
    1 saves
    8 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @altryneChris Alexiuk (@llm_wizard): the quiet cost secret for agents is KV cache. DeepSeek's compression path - down to ~890 bytes per token - is how you keep more context hot without lighting the GPU on fire. Cache hits vs misses is most of the bill.

    3 Sources

    @altryneChris Alexiuk (@llm_wizard): the quiet cost secret for agents is KV cache. DeepSeek's compression path - down to ~890 bytes per token - is how you keep more context hot without lighting the GPU on fire. Cache hits vs misses is most of the bill.