Report
MiMo 2.6 Flash reportedly runs across an RTX 6000 GPU and M5 laptop at 40 tokens per second
The user says llama.cpp supports native mxfp4 weights out of the box; another says ggml RPC distributes inference across devices.
TLDR
A user reports running MiMo 2.6 Flash across an RTX 6000 GPU and an M5 laptop over 10 GbE at 40 tokens per second. They say llama.cpp supports the model’s native mxfp4 weights out of the box. Another user says llama.cpp’s ggml RPC backend distributes inference across different devices, though they describe it as an advanced setting they hope to make more accessible.
Combined views
11K
3 Sources, first seen ago
