Hugging Face Adds Activation Checkpoint Offload to Transformers
Feature reduces GPU memory use for long sequences or larger models.
TLDR
Stas Bekman posted that activation checkpoint offload was merged into the Hugging Face Transformers library. The update targets gradient checkpointing, which had kept one activation per checkpointed layer on the device. At long sequence lengths that term dominates memory use. Bekman stated the change supports longer contexts or bigger models on fewer GPUs without any parallelism. The linked pull request explains the prior limitation and the new offload behavior. A separate PyTorch issue requests matching offload support for torch.util.checkpoint to cpu or nvme.
Combined views
1.5K
1 Source, first seen 29d ago