• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Cache-to-Cache reportedly lets AI models communicate without generating words

    A post says Chinese researchers open-sourced C2C, an approach that passes cached internal data between models instead of text messages. It claims a 2.5× overall speedup.

    SupermanSU
    1 Source, 20d ago, first seen 20d ago

    TLDR

    A post describes Cache-to-Cache (C2C) as an open-source method for language models to communicate through cached internal data, known as the KV cache, rather than generated words. It says a neural network transfers and combines the caches, while a learned gate selects which model layers benefit most. The post claims the approach eliminates intermediate text-generation latency, improves benchmark accuracy by up to 14.2% over individual models and delivers a 2.5× overall speedup.

    Combined views

    150.5K

    1 Source, first seen 20d ago

    Combined views

    150.5K

    1 Source, first seen 20d ago

    3.1K likes
    3.1K likes
    238 comments
    1.7K saves
    596 reposts
    238 comments
    1.7K saves
    596 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Superman@thesupermannxLLMs can now talk to each other without words. Chinese researchers open-sourced a new paradigm that lets LLMs communicate without generating a single word. It’s called Cache-to-Cache (C2C) communication. right now, when multiple ai agents work together, they are forced to translate their internal "thoughts" into human text tokens just to pass a message. this loses rich semantic meaning and causes massive token-by-token latency. So, instead of spitting out words, c2c uses a neural network to directly project and fuse the source model's "kv-cache" right into the target model. it is pure, direct semantic communication.. they even added a learnable gating mechanism to select exactly which layers benefit most from the cache transfer. the benchmark results are actually crazy: - avoids all intermediate text generation latency - accuracy jumps by up to 14.2% compared to individual models - beats traditional text-based agent communication by over 5% - delivers a massive 2.5x speedup in overall speed we are literally watching llms bypass human language to build their own silent, high-speed neural network..20d

    1 Source

    Superman@thesupermannxLLMs can now talk to each other without words. Chinese researchers open-sourced a new paradigm that lets LLMs communicate without generating a single word. It’s called Cache-to-Cache (C2C) communication. right now, when multiple ai agents work together, they are forced to translate their internal "thoughts" into human text tokens just to pass a message. this loses rich semantic meaning and causes massive token-by-token latency. So, instead of spitting out words, c2c uses a neural network to directly project and fuse the source model's "kv-cache" right into the target model. it is pure, direct semantic communication.. they even added a learnable gating mechanism to select exactly which layers benefit most from the cache transfer. the benchmark results are actually crazy: - avoids all intermediate text generation latency - accuracy jumps by up to 14.2% compared to individual models - beats traditional text-based agent communication by over 5% - delivers a massive 2.5x speedup in overall speed we are literally watching llms bypass human language to build their own silent, high-speed neural network..20d