Omni-world models for physical-world AI
Reka Labs has shared research on unifying vision-language-action and world-model approaches under a single backbone.
TLDR
Reka Labs contrasts vision-language models, which take in language and images and output language, with omni models that take in and produce language, images, video and actions. Its research makes the case for combining vision-language-action and world-model approaches under one backbone for physical-world intelligence.
Combined views
1.4K
1 Source, first seen 13h ago
Omni-world models for physical-world AI
Reka Labs has shared research on unifying vision-language-action and world-model approaches under a single backbone.
TLDR
Reka Labs contrasts vision-language models, which take in language and images and output language, with omni models that take in and produce language, images, video and actions. Its research makes the case for combining vision-language-action and world-model approaches under one backbone for physical-world intelligence.