• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Report

    MiMo 2.6 Flash reportedly runs across an RTX 6000 GPU and M5 laptop at 40 tokens per second

    The user says llama.cpp supports native mxfp4 weights out of the box; another says ggml RPC distributes inference across devices.

    Georgi GerganovGG
    merveME
    Pedro CuencaPC
    3 Sources, ,

    TLDR

    A user reports running MiMo 2.6 Flash across an RTX 6000 GPU and an M5 laptop over 10 GbE at 40 tokens per second. They say llama.cpp supports the model’s native mxfp4 weights out of the box. Another user says llama.cpp’s ggml RPC backend distributes inference across different devices, though they describe it as an advanced setting they hope to make more accessible.

    Combined views

    11K

    3 Sources, first seen 6h ago

    Combined views

    11K

    3 Sources, first seen 6h ago

    167 likes
    6h ago
    first seen 6h ago
    167 likes
    10 comments
    50 saves
    30 reposts
    Featured Source
    10 comments
    50 saves
    30 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    Pedro Cuenca@pcuenqIt's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the native mxfp4 weights of a state-of-the-art model, on heterogeneous hardware. Supported out of the box in llama.cpp.6h
    merve@mervenoyannRT @pcuenq: It's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the…1h
    Georgi Gerganov@ggerganovllama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend It's an advanced setting but I think with time we'll make it more accessible to regular users.1h

    3 Sources

    Pedro Cuenca@pcuenqIt's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the native mxfp4 weights of a state-of-the-art model, on heterogeneous hardware. Supported out of the box in llama.cpp.6h
    merve@mervenoyannRT @pcuenq: It's crazy that I can run MiMo 2.6 Flash across my RTX 6000 GPU and my M5 laptop at 40 tokens/sec over 10 GbE 🤯 These are the…1h
    Georgi Gerganov@ggerganovllama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend It's an advanced setting but I think with time we'll make it more accessible to regular users.1h