Report
Meta's Memory Mosaics model is claimed to beat a transformer trained on eight times as many tokens
A post describing a Meta paper says the model uses a network of associative memories instead of standard attention.
TLDR
A post says Meta published “Memory Mosaics at scale,” a paper about a model that uses associative memories rather than standard attention. The post claims a model trained on 1 trillion tokens outperformed a traditional transformer trained on 8 trillion tokens. It also claims stronger in-context learning and the ability to solve new tasks with fewer examples.
Combined views
20.1K
2 Sources, first seen ago
