Looped-transformer decoding reportedly raises AIME 2024 accuracy from 61.9% to 73.3%
A post describing the paper says its method uses the difference between early and final loops to adjust predictions without extra training.
TLDR
A post describing “Decoding Looped Transformers Better for (Almost) Free” says the method treats an early loop as a rough draft and the final loop as a refined answer, then uses the difference to adjust the final prediction. The post reports that Ouro-2.6B-Thinking’s AIME 2024 accuracy rose from 61.9% to 73.3%. It also says using half the recurrent loops can match or beat full-depth decoding while cutting forward FLOPs by up to 48.2%.
Combined views
25.4K
2 Sources, first seen ago
