Reaction
Which open-weight mixture-of-experts coding LLM fits in less than 60GB of RAM?
A user wants something faster than 12 tokens per second and thinks a mixture-of-experts model might help.
TLDR
A user is looking for an open-weight mixture-of-experts LLM for coding that fits in less than 60GB of RAM. They think that type of model might be necessary for reasonably interactive speeds on their hardware and want something faster than 12 tokens per second.
Combined views
18.7K
2 Sources, first seen ago