• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Dust is claimed to approach or sometimes exceed backprop in transformer pretraining with heavy compute

    Dust’s team says it perturbs activations in parallel and is roughly 1,000 to 10,000 times more compute-efficient than EGGROLL for training transformers.

    Cheng LouCL
    SamipSA
    2 Sources, ,

    TLDR

    The team introducing Dust describes it as a zeroth-order method for pretraining transformers. They claim it can approach or sometimes exceed backprop with large amounts of computation. Its core idea is a “virtual population” that perturbs activations in parallel; the team says backprop-like gradients emerge from large populations of those perturbations, with alignment holding at every scale they tested, up to 1 billion tokens.

    Combined views

    36.6K

    2 Sources, first seen 2h ago

    Combined views

    36.6K

    2 Sources, first seen 2h ago

    650 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2h ago
    first seen 2h ago
    650 likes
    15 comments
    840 saves
    58 reposts
    15 comments
    840 saves
    58 reposts

    2 Sources

    Samip@industriaalistBackprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!) - Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. - We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization. - Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones. - Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling. The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel. w/ @bishmdl76, @cs_serdar, @akshayvegesna2h
    Cheng Lou@_chenglouRT @industriaalist: Backprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the fi…2h
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    Samip@industriaalistBackprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the first zeroth-order method to pretrain transformers to approach and sometimes even *exceed* backprop with large amounts of computation (bitter lesson!) - Dust is on the order of 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. - We chose pretraining transformers because it's arguably the hardest task possible for zeroth-order optimization. - Search in high dimensions is really misunderstood: zeroth-order usually gets better with overparameterization, not worse, i.e. at a fixed population size, larger models (up to 120x) reach a lower loss than smaller ones. - Backprop-like gradients emerge from a large population of activation perturbations, and the alignment holds up at every scale we tested, up to 1B tokens, which is critical for scaling. The core idea is a *virtual population*, which helps scale to large populations in transformers by perturbing activations in parallel. w/ @bishmdl76, @cs_serdar, @akshayvegesna2h
    Cheng Lou@_chenglouRT @industriaalist: Backprop has been the only credit assignment algorithm capable of training large neural nets. Introducing Dust, the fi…2h