Announcement
Local AI slide deck covers prefill, decode and quantization with llama.cpp
The creator says it also covers MoE versus dense models and VRAM versus unified memory, and welcomes reuse with attribution.
TLDR
A user says they released a local AI slide deck covering prefill versus decode, MoE versus dense models, VRAM versus unified memory, quantization and speculative decoding, all with llama.cpp. They invite others to reuse it with attribution.
Combined views
7.8K
2 Sources, first seen ago
123 likes17 comments100 saves23 reposts
