Announcement
Could looped language models match Transformers with 3× fewer parameters and a 3× smaller KV cache?
A post introducing “Rethinking at Fixed Points” proposes fixed-point shortcuts for training and inference.
TLDR
A post introducing “Rethinking at Fixed Points” asks whether looped language models could use 3× fewer parameters and a 3× smaller KV cache while remaining comparable to a standard Transformer. It argues that adding computing work could let the models scale without increasing memory use, and proposes fixed-point shortcuts, a learned depth prior and orthogonal injection.
Combined views
7.4K
2 Sources, first seen ago
