Patch Policy Feeds Frozen Pretrained ViT Tokens Into Policy Models
New method feeds frozen pretrained ViT patch tokens into transformer policy heads for robotics tasks.
Patch Policy is a new robotics approach that extracts patch tokens from an internet-pretrained frozen Vision Transformer and feeds them into models such as VQ-BeT or Diffusion Policy. Researchers report it delivers massive gains in training efficiency, sharply cuts model size and latency, and outperforms OpenVLA-OFT while using only 0.7 percent of the parameters. The method was highlighted by academics Furong Huang and Lerrel Pinto as a simple switch from standard visual tokens to dense features. Status remains at the research announcement stage with no deployment details confirmed.
Combined views
4K
2 posts, first seen 18h ago