• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    SemiAnalysis describes speed trade-offs in positional embeddings

    SemiAnalysis says a lookup-table approach performed fine in a toy test but took around five times longer to test than earlier approaches in a naïve MacBook implementation.

    SE
    7 Sources, 18d ago, first seen 18d ago

    TLDR

    SemiAnalysis compares approaches to positional embeddings—ways of representing token positions. One uses a lookup table that maps relative displacement to a learnable matrix. The publisher says this preserves translation invariance, but calls it a poor use of computation and warns it would be brittle in practice. Another approach learns separate embedding functions for the two token positions. SemiAnalysis says it does not preserve translation invariance, but could be useful if information were encoded in individual positions rather than relative displacement. Despite having more learnable parameters, that approach ran much faster than the lookup table after a simple optimization, according to SemiAnalysis.

    Combined views

    89.8K

    7 Sources, first seen 18d ago

    Combined views

    89.8K

    7 Sources, first seen 18d ago

    606 likes
    606 likes
    17 comments
    564 saves
    39 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    17 comments
    564 saves
    39 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 Sources

    @SemiAnalysis_Starting with the hard part, let's write down Jane Street's math. Alok Puranik explores positional embeddings that allow the attention score to be written q(s)T F(s)T G(t) k(t), where q and k are the position-free queries and keys, and F and G do the embedding work. This encodes the assumption that the positional embedding must act linearly and separably on q and k. Puranik then folds F(s)T G(t) into a single A(t-s), stipulating that the embedding is only sensible if it is translation invariant. The last requirement he makes is that A(0) = I, which can be satisfied without loss of generality as long as A is nondegenerate. With a bit of algebra, he derives the group law, and, assuming continuity, shows that embedding function must have the form exp((t-s)X) for some generator matrix X. The study of all positional embeddings gets reduced to the study of the these generator matrices—a surprisingly strong result for relatively weak assumptions. (2/7)

    7 Sources

    @SemiAnalysis_Starting with the hard part, let's write down Jane Street's math. Alok Puranik explores positional embeddings that allow the attention score to be written q(s)T F(s)T G(t) k(t), where q and k are the position-free queries and keys, and F and G do the embedding work. This encodes the assumption that the positional embedding must act linearly and separably on q and k. Puranik then folds F(s)T G(t) into a single A(t-s), stipulating that the embedding is only sensible if it is translation invariant. The last requirement he makes is that A(0) = I, which can be satisfied without loss of generality as long as A is nondegenerate. With a bit of algebra, he derives the group law, and, assuming continuity, shows that embedding function must have the form exp((t-s)X) for some generator matrix X. The study of all positional embeddings gets reduced to the study of the these generator matrices—a surprisingly strong result for relatively weak assumptions. (2/7)