Report
Cerebras' wafer-scale processor is claimed to move model weights 2,500 times faster than a GPU during LLM decoding
A post sharing Cerebras CEO Andrew Feldman's explanation credits the claimed speed gap to SRAM rather than GPUs' HBM.
TLDR
A post sharing Cerebras CEO Andrew Feldman's podcast discussion describes LLM inference as prompt processing followed by generating tokens one at a time. It says Cerebras keeps model weights in SRAM across its wafer-scale processor, while GPUs fetch them from HBM, and claims moving those weights into compute during decoding is about 2,500 times faster.
Combined views
7.1K
2 Sources, first seen ago
