• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Stanford Prefix Sliding Paper Speeds Reasoning Inference

    Rohan Paul shares arXiv screenshot of Stanford paper on efficient long reasoning.

    RP
    2 Sources, 27d ago, first seen 27d ago

    TLDR

    Rohan Paul tweeted about a new Stanford paper titled Prefix Sliding. According to the post, long reasoning does not require the full chain of thought in memory. Retaining only the task prefix and recent tokens allows inference to run about three times faster without any retraining. The method targets the expense of full attention over growing token sequences. The tweet includes an attached screenshot of the arXiv paper itself.

    Combined views

    5.4K

    2 Sources, first seen 27d ago

    Combined views

    5.4K

    2 Sources, first seen 27d ago

    39 likes
    39 likes
    7 comments
    29 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 comments
    29 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @rohanpaul_aiNew Stanford paper Prefix Sliding shows that long reasoning does not need the full chain of thought in memory: keep the task prefix and recent tokens, and inference can run about 3x faster without retraining. Long reasoning gets expensive because full attention makes every new token look back over an ever-growing chain of thought. Prefix Sliding keeps the fixed prefix with task and tool instructions plus a sliding window of recent reasoning, dropping older middle tokens as it goes. Their attention analysis points in the same direction: the prefix and latest tokens receive most attention, while intermediate reasoning gets little. On Qwen3-1.7B, a 4,096-token window scored 33.9% on AIME25 versus 34.2% with full attention, and the paper reports about 3x faster inference without retraining. Once the window fills, each new token has constant attention cost instead of becoming more expensive with every step, which also enabled reinforcement-learning rollouts beyond 100,000 tokens. Prefix Sliding beat pure sliding windows, repeated summarization, and last-k deletion on the tested speed/accuracy tradeoff. – arxiv. org/abs/2608.26070 Title: "Prefix Sliding for efficient test-time scaling"

    2 Sources

    @rohanpaul_aiNew Stanford paper Prefix Sliding shows that long reasoning does not need the full chain of thought in memory: keep the task prefix and recent tokens, and inference can run about 3x faster without retraining. Long reasoning gets expensive because full attention makes every new token look back over an ever-growing chain of thought. Prefix Sliding keeps the fixed prefix with task and tool instructions plus a sliding window of recent reasoning, dropping older middle tokens as it goes. Their attention analysis points in the same direction: the prefix and latest tokens receive most attention, while intermediate reasoning gets little. On Qwen3-1.7B, a 4,096-token window scored 33.9% on AIME25 versus 34.2% with full attention, and the paper reports about 3x faster inference without retraining. Once the window fills, each new token has constant attention cost instead of becoming more expensive with every step, which also enabled reinforcement-learning rollouts beyond 100,000 tokens. Prefix Sliding beat pure sliding windows, repeated summarization, and last-k deletion on the tested speed/accuracy tradeoff. – arxiv. org/abs/2608.26070 Title: "Prefix Sliding for efficient test-time scaling"