• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GLM 5.2 1.6TB Policy Transferred in About 9 Seconds

    Engineer highlights trainer-to-inference weight transfer in prime-rl for large models.

    EL
    PI
    SA
    12 Sources, 27d ago, first seen 27d ago

    TLDR

    Lan Dao posted about work by m_sirovatka and team on prime-rl. The update supports trainer-to-inference transfer of the full 1.6TB policy for GLM 5.2. Reported times reach roughly 9 seconds, with some experiments completing in under 4 seconds. m_sirovatka described the effort as a solution for blazingly fast weight transfer after focused development time. The posts present the capability as an improvement that reduces step time.

    Combined views

    37.6K

    12 Sources, first seen 27d ago

    Combined views

    37.6K

    12 Sources, first seen 27d ago

    407 likes
    407 likes
    20 comments
    127 saves
    62 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    20 comments
    127 saves
    62 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    12 Sources

    @PrimeIntellectOur RL stack now supports NIXL weight transfer, reducing trainer-to-inference transfer time 9x compared with NCCL: from 86 seconds down to single-digit seconds for an 800B-parameter model, and even <4 seconds in our experiments. For prime-rl users, this means over 25% more throughput end-to-end compared with our previous speed. It also clears the way for fault-tolerant, elastic inference scaling that NCCL's rigid process groups made difficult.
    @m_sirovatkayou can do blazingly fast weight transfer in prime-rl now, spent quite a while on this (imo quite cool) solution and finally happy with the performance, give it a try
    @ad0rnaiExcellent work by @m_sirovatka and team on further reducing step time, this time with trainer-to-inference transfer Now you can transfer the full 1.6TB policy of GLM 5.2 in ~9 seconds and even <4 seconds in some experiments
    @eliebakouchRT @PrimeIntellect: Our RL stack now supports NIXL weight transfer, reducing trainer-to-inference transfer time 9x compared with NCCL: from…
    @KranenKyleFast refit in RL runs allows you to cut down on dead GPU time during which weights are being loaded. Our work on NIXL and MX with the Prime Intellect team demonstrated a reduction of refit time from 86 seconds to 4 seconds: https://www.primeintellect.ai/blog/nixl-modelexpress-weight-transfer
    @samsja19RT @KranenKyle: Fast refit in RL runs allows you to cut down on dead GPU time during which weights are being loaded. Our work on NIXL and…

    12 Sources

    @PrimeIntellectOur RL stack now supports NIXL weight transfer, reducing trainer-to-inference transfer time 9x compared with NCCL: from 86 seconds down to single-digit seconds for an 800B-parameter model, and even <4 seconds in our experiments. For prime-rl users, this means over 25% more throughput end-to-end compared with our previous speed. It also clears the way for fault-tolerant, elastic inference scaling that NCCL's rigid process groups made difficult.
    @m_sirovatkayou can do blazingly fast weight transfer in prime-rl now, spent quite a while on this (imo quite cool) solution and finally happy with the performance, give it a try
    @ad0rnaiExcellent work by @m_sirovatka and team on further reducing step time, this time with trainer-to-inference transfer Now you can transfer the full 1.6TB policy of GLM 5.2 in ~9 seconds and even <4 seconds in some experiments
    @eliebakouchRT @PrimeIntellect: Our RL stack now supports NIXL weight transfer, reducing trainer-to-inference transfer time 9x compared with NCCL: from…
    @KranenKyleFast refit in RL runs allows you to cut down on dead GPU time during which weights are being loaded. Our work on NIXL and MX with the Prime Intellect team demonstrated a reduction of refit time from 86 seconds to 4 seconds: https://www.primeintellect.ai/blog/nixl-modelexpress-weight-transfer
    @samsja19RT @KranenKyle: Fast refit in RL runs allows you to cut down on dead GPU time during which weights are being loaded. Our work on NIXL and…