A reader recommends a free book on training large language models at scale
After reading the free online version on Hugging Face, the reader says they bought a physical copy for their library.
TLDR
A reader calls the technical book “arguably one of the best” for understanding how large language models are trained at scale. They highlight GPU memory and profiling; tiling, kernel fusion and FlashAttention; and data, tensor, pipeline and context parallelism. They share the free online version here: https://huggingface.co/spaces/nanotron/ultrascale-playbook
Combined views
65.7K
1 Source, first seen 24d ago