• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Standard looping is reportedly compute-optimal for multi-epoch training

    For multi-epoch training, a project participant says the optimal number of loops increases with the compute budget and looping has a useful regularizing effect.

    Andrew Gordon WilsonAG
    Songlin YangSY
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)T(
    6 Sources, ,

    TLDR

    A project participant reports that even standard looping is compute-optimal for multi-epoch training—training over repeated passes through the data. They say the optimal number of loops increases with the computational budget and that looping has a useful regularizing effect in that setting. They also identify effective depth as a key factor in both single- and multi-epoch training.

    Combined views

    74.9K

    6 Sources, first seen 20d ago

    Combined views

    74.9K

    6 Sources, first seen 20d ago

    507 likes
    20d ago
    first seen 20d ago
    507 likes
    18 comments
    417 saves
    82 reposts
    18 comments
    417 saves
    82 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    Samip@industriaalistWe've been obsessed with bending the scaling laws recently and our new paper shows that looping with model growth improves the scaling exponent, leading to gain that compounds with compute! Everyone assumes architectural changes only give constant factor gains and pretraining progress mostly comes from data (e.g. @dwarkesh_sp's recent post). We found that model growth, looping, and boundary operators result in compute multipliers over standard transformers that grow exponentially with each OOM of compute. - 1.55x at 1e20 FLOPs and 2.7x projected at 1e25. - Matches GPT-3 13B on CORE with 20x less compute w/ @charllechen, @akshayvegesna, @andrewgwils 🧵20d
    Andrew Gordon Wilson@andrewgwilsMuch more in the paper! This was an exciting project, led by @charllechen. Please also see Charlie's thread: https://x.com/charllechen/status/2100601796011995437 and Samip's thread: 7/720d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexMakes sense I don't know why anyone would assume that architecture only offers constant gains. And the effective depth perspective is the most salient candidate for improvement here. I think this is what frontier labs are working on, what determines willingness to go for 5-10-20T20d
    kalomaze@kalomaze@teortaxesTex >especially as we become data constrained mmm i don't like the opinion (that being data constrained is an inherent irreducible property of the modeling-general-data regime, or hell, even narrowly wrt "information for niche domains") being laundered into a binding constraint here20d
    Songlin Yang@SonglinYang4RT @industriaalist: We've been obsessed with bending the scaling laws recently and our new paper shows that looping with model growth impro…20d

    6 Sources

    Samip@industriaalistWe've been obsessed with bending the scaling laws recently and our new paper shows that looping with model growth improves the scaling exponent, leading to gain that compounds with compute! Everyone assumes architectural changes only give constant factor gains and pretraining progress mostly comes from data (e.g. @dwarkesh_sp's recent post). We found that model growth, looping, and boundary operators result in compute multipliers over standard transformers that grow exponentially with each OOM of compute. - 1.55x at 1e20 FLOPs and 2.7x projected at 1e25. - Matches GPT-3 13B on CORE with 20x less compute w/ @charllechen, @akshayvegesna, @andrewgwils 🧵20d
    Andrew Gordon Wilson@andrewgwilsMuch more in the paper! This was an exciting project, led by @charllechen. Please also see Charlie's thread: https://x.com/charllechen/status/2100601796011995437 and Samip's thread: 7/720d
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)@teortaxesTexMakes sense I don't know why anyone would assume that architecture only offers constant gains. And the effective depth perspective is the most salient candidate for improvement here. I think this is what frontier labs are working on, what determines willingness to go for 5-10-20T20d
    kalomaze@kalomaze@teortaxesTex >especially as we become data constrained mmm i don't like the opinion (that being data constrained is an inherent irreducible property of the modeling-general-data regime, or hell, even narrowly wrt "information for niche domains") being laundered into a binding constraint here20d
    Songlin Yang@SonglinYang4RT @industriaalist: We've been obsessed with bending the scaling laws recently and our new paper shows that looping with model growth impro…20d