Report
MiMo 2.6 Flash reportedly runs across an RTX 6000 GPU and M5 laptop at 40 tokens per second
A user says the setup uses native mxfp4 weights and works out of the box in llama.cpp.
TLDR
A user reports running MiMo 2.6 Flash across an RTX 6000 GPU and M5 laptop at 40 tokens per second over 10 GbE. They say the run uses native mxfp4 weights on the two types of hardware and is supported out of the box in llama.cpp.
Combined views
7.6K
2 Sources, first seen ago
50 likes5 comments18 saves23 reposts
