Reaction
Multimodal LLM is claimed to tackle a robotics task it was never deliberately trained for
The demonstrator says DSV4.1F uses unfamiliar cameras and motors; the video runs at 10x speed.
TLDR
A demonstrator says DSV4.1F, a multimodal LLM, tackles a task it was never deliberately trained for using cameras it has never seen and motors it did not learn on. The video runs at 10x speed because the system is currently slow. The demonstrator argues that general-purpose models may displace much robotics task training, while reinforcement learning remains valuable for low-level skills like gait and reaching.
Combined views
4.4K
2 Sources, first seen ago
