Next-token prediction and the search for more data-efficient robotics
A user hopes for robotics training that needs less data, but doubts that a better training objective will look very different from next-token prediction.
TLDR
A post questions how far better training objectives might depart from next-token prediction. The author describes multi-token prediction, offline reinforcement learning and inverse reinforcement learning as weighted variants of predicting future tokens. In their view, nothing substantially different has convincingly shown a step change or offered both low computational cost and robustness when scaling. They still hope for a more data-efficient approach for robotics.
Combined views
1.2K
2 Sources, first seen 16d ago