Visual AI as a possible path to AGI
The post describes a paper involving Google DeepMind, Harvard, Stanford and other labs that calls for visual systems to model the world, remember changes, predict outcomes and act.
TLDR
A post summarizing the paper says vision should do more than supply input to a language model. It describes a possible route toward artificial general intelligence: systems that learn from images, video, 3D structure and interaction, keep updating their understanding of the world, and use it to predict and act. According to the summary, the paper points to persistent memory, continual learning and robotics among the possible pieces—and argues that image Q&A, captions and realistic video should not be the main ways to judge visual AI.
Combined views
23.7K
3 Sources, first seen 19d ago