Danfei Xu Identifies Sim2Real and Behavior Cloning as Core Robot Learning Shifts
Reactions from ranked influencers
6 postsMany groundbreaking works grew off these two roots. But few other trees cross over. VLMs and video models are two notable grafts. More on this later. 2/
But the three forms differ in their scaling curves. Consider y-axis: capability = unique tasks × success rate; x-axis: time × resources. Teleop got a head start, but will grow slowly and plateau at the lowest point. UMI-like may scale up to whatever its gripper/glove form factor allows. Ego + dex hands has the highest ceiling (human-level), but has embodiment and dex-hand hardware gaps to close, so it starts the slowest. 4/
Who & when will get to the “GPT moment”? My definition is: 40% success rate, any task, anywhere. My bet is UMI-like will get there first, given the current investment and technology. Ego will follow soon after. Teleop may never, not because of infeasibility, but because of impatience. 5/
The S2R locomotion tree has reached maturity. Perhaps more growth is possible. The BC manipulation tree is still a sapling. Whatever the form, teleop, UMI-like, or ego, the fundamental bet is that we can BC human behaviors. 3/
What about VLA/WAM, or whatever next prior-rich architecture? My best guess is that they change the slope of these curves, but not the asymptotic ceiling. They may get there faster, but at sufficient scale the ceilings may become indistinguishable. 6/
When will we get to ROI? Attempts in RL/DAgger post-training continuously happen at various points on this curve. Specialization has value. But I fear the current crazed investment is pushing ROI later in the timeline. I hope it comes before this round of the bubble bursts. 7/
Combined views
16.3K
6 posts, first seen 14h ago