• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Shared-cache transformer training tweak reportedly removes 81–89% of the primary continuation penalty

    The researcher says generation had higher loss than prefill, even with ground-truth token inputs.

    AK
    5 Sources, 1h ago, first seen 1h ago

    TLDR

    After leaving Google DeepMind, a researcher describes a short, compute-constrained project on a transformer with one shared core key-value cache. They say their two-pass training still produced cache histories that differed from those during continuation. Matching the passes’ donor keys and values during the final 10% of training removed 81–89% of the primary continuation penalty at three tested widths, they report. The token budget stayed the same, while full-cache loss rose by 0.0011–0.0025 nats.

    Combined views

    2.6K

    5 Sources, first seen 1h ago

    Combined views

    2.6K

    5 Sources, first seen 1h ago

    39 likes
    Architecture of the default tied transformer: one untied prologue block, four distinct core blocks reused for three iterations with an LSTM-style carry, then two untied epilogue blocks. The core reads one common donor KV cache for preceding tokens, produced by its final block on its final iteration. Each block retains its own current-token KV. There are seven block weight sets and fifteen block applications. With separate prologue and epilogue caches, sharing reduces persistent cache sets from fifteen to four.
    39 likes
    6 comments
    17 saves
    4 reposts
    6 comments
    17 saves
    4 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @BlackHCA first (short & compute-constrained) research project after leaving Google DeepMind: an LSTM-style tied transformer with one shared core KV cache The main question is how to train for the cache histories that the model produces recursively at inference1h

    5 Sources

    @BlackHCA first (short & compute-constrained) research project after leaving Google DeepMind: an LSTM-style tied transformer with one shared core KV cache The main question is how to train for the cache histories that the model produces recursively at inference1h