The prefill-versus-decode split in AI compute
A post sharing Chamath’s comments describes prefill as compute-bound and decode as memory-bandwidth-bound.
TLDR
The post says massively parallel GPUs give Nvidia an advantage in prefill as context grows. It describes decode as constrained by memory bandwidth because each new token depends on scanning what has already been generated.
Combined views
3.4K
1 Source, first seen 6h ago
19 likes