• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

A proposed learning method uses a growing gradient-signal cache instead of per-step linear-layer updates

A researcher says the paper derives the method from a duality between SGD and linear attention.

Han GuoHG
Oliver Sieberling @ COLMOS
2 Sources, 6h ago, first seen 6h ago

TLDR

A researcher says a new paper with a Google PI proposes a learning method derived from the duality between SGD and linear attention. Instead of updating linear layers with gradient descent at every training step, the method maintains a growing “KV cache” of gradient signals to attend over. The researcher says the neural network grows during training.

Combined views

2.4K

2 Sources, first seen 6h ago

65 likes42 saves19 reposts

Combined views

2.4K

2 Sources, first seen 6h ago

65 likes42 saves19 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

2 Sources

Oliver Sieberling @ COLM@osieberlingNew paper with Google PI: We derive a new learning paradigm coming from the sgd/linear-attention duality, where instead of updating linear layers with gd at every step during training, we maintain a growing “kv cache” of gradient signals we attend over (NN grows during training)6h
Han Guo@HanGuo97RT @osieberling: New paper with Google PI: We derive a new learning paradigm coming from the sgd/linear-attention duality, where instead of…19m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    2 Sources

    Oliver Sieberling @ COLM@osieberlingNew paper with Google PI: We derive a new learning paradigm coming from the sgd/linear-attention duality, where instead of updating linear layers with gd at every step during training, we maintain a growing “kv cache” of gradient signals we attend over (NN grows during training)6h
    Han Guo@HanGuo97RT @osieberling: New paper with Google PI: We derive a new learning paradigm coming from the sgd/linear-attention duality, where instead of…19m
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet