Dust is claimed to approach or sometimes exceed backprop in transformer pretraining with heavy compute
Dust’s team says it perturbs activations in parallel and is roughly 1,000 to 10,000 times more compute-efficient than EGGROLL for training transformers.
TLDR
The team introducing Dust describes it as a zeroth-order method for pretraining transformers. They claim it can approach or sometimes exceed backprop with large amounts of computation. Its core idea is a “virtual population” that perturbs activations in parallel; the team says backprop-like gradients emerge from large populations of those perturbations, with alignment holding at every scale they tested, up to 1 billion tokens.
Combined views
36.6K
2 Sources, first seen ago
