• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Hugging Face training loop reportedly reaches the same reward in 53 minutes, down from 3 hours 27

    A user describes a reinforcement learning setup that keeps training, model serving and storage on Hugging Face, syncing only small adapter updates through Buckets.

    Lewis Tunstall @ COLM 🌉LT
    Adithya S KAS
    2 Sources, ,

    TLDR

    A user highlights a reinforcement learning loop that trains on one Hugging Face Job and uses vLLM to serve the model on others. The setup uses asynchronous GRPO with LoRA in TRL, syncing only the small adapter update through a Bucket. The user reports that the optimized loop takes 53 minutes instead of 3 hours 27 minutes for the same reward.

    Combined views

    16.6K

    2 Sources, first seen 20d ago

    Combined views

    16.6K

    2 Sources, first seen 20d ago

    267 likes
    20d ago
    first seen 20d ago
    267 likes
    9 comments
    186 saves
    27 reposts
    9 comments
    186 saves
    27 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    Adithya S K@adithya_s_kYou can now do RL weight sync over HF Buckets. > Async GRPO + LoRA in TRL > train on one Job, serve vLLM on others > sync just the small delta adapter through a Bucket. Trainer, inference and storage, all on @huggingface infra. Beautifully optimised loop by @DirhoussssiAmine 👇 3h27 → 53 min for the same reward.20d
    Lewis Tunstall @ COLM 🌉@_lewtunRT @adithya_s_k: You can now do RL weight sync over HF Buckets. > Async GRPO + LoRA in TRL > train on one Job, serve vLLM on others > sync…20d

    2 Sources

    Adithya S K@adithya_s_kYou can now do RL weight sync over HF Buckets. > Async GRPO + LoRA in TRL > train on one Job, serve vLLM on others > sync just the small delta adapter through a Bucket. Trainer, inference and storage, all on @huggingface infra. Beautifully optimised loop by @DirhoussssiAmine 👇 3h27 → 53 min for the same reward.20d
    Lewis Tunstall @ COLM 🌉@_lewtunRT @adithya_s_k: You can now do RL weight sync over HF Buckets. > Async GRPO + LoRA in TRL > train on one Job, serve vLLM on others > sync…20d