• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    “Free pause tokens” reportedly add AI compute without longer context

    DAIR.AI says a paper from Microsoft and Cornell describes essentially no extra inference latency and an unchanged memory cache, with the extra cost shifted to training.

    You JiachengYJ
    Liliang RenLR
    2 Sources, ,

    TLDR

    DAIR.AI describes “free pause tokens” as extra computation carried in a parallel prediction stream that shares the model’s weights, rather than taking up additional sequence positions. It says the approach leaves context length and the KV cache unchanged, with essentially no extra latency during inference. DAIR.AI puts training overhead at about 1.14× an optimized pretraining pipeline while retaining most of the benefit. A quote-post questions the idea’s novelty, saying people “keep reinventing ideas from XLNet.”

    Combined views

    33.5K

    2 Sources, first seen 22d ago

    Combined views

    33.5K

    2 Sources, first seen 22d ago

    139 likes
    22d ago
    first seen 22d ago
    139 likes
    6 comments
    88 saves
    12 reposts
    6 comments
    88 saves
    12 reposts

    2 Sources

    Liliang Ren@liliang_renPeople keep reinventing ideas from XLNet smh https://arxiv.org/abs/1906.0823722d
    You Jiacheng@YouJiachengshould we introduce two-stream for CED? Causal Encoder encode BOTH context (for later tokens) and query (for self) in a single stream. It may even partially decode the token.21d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Liliang Ren@liliang_renPeople keep reinventing ideas from XLNet smh https://arxiv.org/abs/1906.0823722d
    You Jiacheng@YouJiachengshould we introduce two-stream for CED? Causal Encoder encode BOTH context (for later tokens) and query (for self) in a single stream. It may even partially decode the token.21d