A claimed way to pretrain transformers without backpropagation
A poster says they've figured out how to use zeroth-order optimization for transformer pretraining and argues that many core assumptions in optimization research are wrong.
TLDR
A September 28 post claims its author has figured out how to pretrain transformers with zeroth-order optimization and no backpropagation. The author said a paper would be out soon.
Combined views
129K
3 Sources, first seen 4h ago
1.5K likes