Announcement
Fast Metal kernels in llama.cpp now cover all 26 Mac GPU weight formats
The contributor says the latest PR added 16 formats, including BF16 and MXFP4, to the 10 covered earlier.
TLDR
The contributor says a newly merged llama.cpp PR extends fast Metal kernels from 10 Mac GPU weight formats to all 26, including BF16, MXFP4, Q2_K, Q3_K and IQ quants. They say the earlier kernels sped up speculative decoding for 10 formats, while matrix multiplies on an M3 Ultra run up to 4.4 times faster when a model checks several tokens at once.
Combined views
4K
2 Sources, first seen ago
