• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    Looped-transformer decoding reportedly raises AIME 2024 accuracy from 61.9% to 73.3%

    A post describing the paper says its method uses the difference between early and final loops to adjust predictions without extra training.

    KK
    AL
    2 Sources, ,

    TLDR

    A post describing “Decoding Looped Transformers Better for (Almost) Free” says the method treats an early loop as a rough draft and the final loop as a refined answer, then uses the difference to adjust the final prediction. The post reports that Ouro-2.6B-Thinking’s AIME 2024 accuracy rose from 61.9% to 73.3%. It also says using half the recurrent loops can match or beat full-depth decoding while cutting forward FLOPs by up to 48.2%.

    Combined views

    25.4K

    2 Sources, first seen 15h ago

    Combined views

    25.4K

    2 Sources, first seen 15h ago

    494 likes
    15h ago
    first seen 15h ago
    494 likes
    10 comments
    363 saves
    113 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    10 comments
    363 saves
    113 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @askalphaxiv"Decoding Looped Transformers Better for (Almost) Free" This paper treats an early loop from a looped transformer as a rough draft and the final loop as the refined answer. It then uses the difference between them to push the final prediction further in the direction the model was already improving, with no extra training. On Ouro-2.6B-Thinking, AIME 2024 accuracy jumps from 61.9% to 73.3% And using only half the recurrent loops can still match or beat full-depth decoding while cutting forward FLOPs by up to 48.2% https://www.alphaxiv.org/abs/2610.0218515h
    @kastnerkyleRT @askalphaxiv: "Decoding Looped Transformers Better for (Almost) Free" This paper treats an early loop from a looped transformer as a ro…5h

    2 Sources

    @askalphaxiv"Decoding Looped Transformers Better for (Almost) Free" This paper treats an early loop from a looped transformer as a rough draft and the final loop as the refined answer. It then uses the difference between them to push the final prediction further in the direction the model was already improving, with no extra training. On Ouro-2.6B-Thinking, AIME 2024 accuracy jumps from 61.9% to 73.3% And using only half the recurrent loops can still match or beat full-depth decoding while cutting forward FLOPs by up to 48.2% https://www.alphaxiv.org/abs/2610.0218515h
    @kastnerkyleRT @askalphaxiv: "Decoding Looped Transformers Better for (Almost) Free" This paper treats an early loop from a looped transformer as a ro…5h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet