How dense and mixture-of-experts models use parameters differently
Nvidia introduces its technical explainer with a question: how can a 30-billion-parameter model activate just 3 billion parameters per token and still draw on its full capacity?
TLDR
Nvidia says its new technical explainer compares how dense and mixture-of-experts (MoE) models use parameters, and what those differences mean for throughput, memory and the complexity of serving models.
Combined views
16.1K
1 Source, first seen 15d ago
171 likes