• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    New Metal kernels for speculative decoding land in llama.cpp

    Qvac says the Apple Silicon update makes speculative decoding up to 3.4 times faster than plain decoding on an M3 Ultra.

    GG
    QV
    2 Sources, ,

    TLDR

    Qvac says its pull request adding Metal kernels for speculative decoding on Apple Silicon was merged into llama.cpp. On an M3 Ultra, it reports 110 tokens per second for speculative decoding versus 32.1 for plain decoding. Qvac says speculative decoding on a Mac was slower than plain decoding before the change.

    Combined views

    3.2K

    2 Sources, first seen 3h ago

    Combined views

    3.2K

    2 Sources, first seen 3h ago

    55 likes
    3h ago
    first seen 3h ago
    55 likes
    7 comments
    7 saves
    8 reposts
    7 comments
    7 saves
    8 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #8

    Today's Rank

    #8

    2 Sources

    @qvacOne of our PR just got merged into llama.cpp: new Metal kernels for speculative decoding on Apple Silicon. Before this PR, speculative decoding on a Mac was slower than plain decoding. On an M3 Ultra it now runs up to 3.4x faster than plain decoding, 110 tok/s against 32.1. Thanks to @ggerganov for reviewing and refining it. http://github.com/ggml-org/llama.cpp/pull/298693h
    @ggerganovRT @qvac: One of our PR just got merged into llama.cpp: new Metal kernels for speculative decoding on Apple Silicon. Before this PR, specu…2h

    2 Sources

    @qvacOne of our PR just got merged into llama.cpp: new Metal kernels for speculative decoding on Apple Silicon. Before this PR, speculative decoding on a Mac was slower than plain decoding. On an M3 Ultra it now runs up to 3.4x faster than plain decoding, 110 tok/s against 32.1. Thanks to @ggerganov for reviewing and refining it. http://github.com/ggml-org/llama.cpp/pull/298693h
    @ggerganovRT @qvac: One of our PR just got merged into llama.cpp: new Metal kernels for speculative decoding on Apple Silicon. Before this PR, specu…2h