Report
The claimed connection between 1991 ULTRA and Google's 2017 Transformer
A post argues that ULTRA's computational costs scale linearly with input size, unlike the later Transformer's quadratic costs.
TLDR
A post argues that Google's 2017 Transformer is based on principles of the 1991 unnormalized linear Transformer, or ULTRA. It says ULTRA's computational costs scale linearly with input size rather than quadratically, and claims a 1993 paper on a recurrent ULTRA extension introduced the terminology of attention.
Combined views
4.5K
4 Sources, first seen ago