• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Hugging Face Adds Activation Checkpoint Offload to Transformers

    Feature reduces GPU memory use for long sequences or larger models.

    SB
    1 Source, 29d ago, first seen 29d ago

    TLDR

    Stas Bekman posted that activation checkpoint offload was merged into the Hugging Face Transformers library. The update targets gradient checkpointing, which had kept one activation per checkpointed layer on the device. At long sequence lengths that term dominates memory use. Bekman stated the change supports longer contexts or bigger models on fewer GPUs without any parallelism. The linked pull request explains the prior limitation and the new offload behavior. A separate PyTorch issue requests matching offload support for torch.util.checkpoint to cpu or nvme.

    Combined views

    1.5K

    1 Source, first seen 29d ago

    Combined views

    1.5K

    1 Source, first seen 29d ago

    22 likes
    22 likes
    11 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 saves
    1 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @StasBekmanActivation checkpoint offload has been just added to @huggingface Transformers. This is super useful for enabling long context or using larger models with less gpus, w/o using any parallelism as it frees up a ton of gpu memory https://github.com/huggingface/transformers/pull/48444 I don't understand why @PyTorch doesn't implement this https://github.com/pytorch/pytorch/issues/158657 since it belongs in the core and it's a power feature almost everybody should use if it's done via a full overlap as it can come at 0-cost to performance.

    1 Source

    @StasBekmanActivation checkpoint offload has been just added to @huggingface Transformers. This is super useful for enabling long context or using larger models with less gpus, w/o using any parallelism as it frees up a ton of gpu memory https://github.com/huggingface/transformers/pull/48444 I don't understand why @PyTorch doesn't implement this https://github.com/pytorch/pytorch/issues/158657 since it belongs in the core and it's a power feature almost everybody should use if it's done via a full overlap as it can come at 0-cost to performance.