Matryoshka LM Suites Train Model Families in One Run
New approach trains multiple model sizes jointly inside one architecture.
TLDR
Nathan Godey posted a thread announcing a paper on Matryoshka LM Suites. The post states that nesting multiple language model sizes inside one architecture permits training the full suite in a single run. Godey claims this yields much less compute while keeping performance equivalent to independent training runs. The same post asserts that the nesting also allows free distillation and better speculative decoding. Yoav Artzi shared the thread. Visible replies came from Yoav Goldberg, who called the work nice, and from Leshem Choshen, who discussed distillation effects on the smallest weights.
Combined views
68.4K
10 Sources, first seen 42d ago
