• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Visual AI as a possible path to AGI

    The post describes a paper involving Google DeepMind, Harvard, Stanford and other labs that calls for visual systems to model the world, remember changes, predict outcomes and act.

    RP
    3 Sources, 19d ago, first seen 19d ago

    TLDR

    A post summarizing the paper says vision should do more than supply input to a language model. It describes a possible route toward artificial general intelligence: systems that learn from images, video, 3D structure and interaction, keep updating their understanding of the world, and use it to predict and act. According to the summary, the paper points to persistent memory, continual learning and robotics among the possible pieces—and argues that image Q&A, captions and realistic video should not be the main ways to judge visual AI.

    Combined views

    23.7K

    3 Sources, first seen 19d ago

    Combined views

    23.7K

    3 Sources, first seen 19d ago

    277 likes
    277 likes
    24 comments
    145 saves
    68 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    24 comments
    145 saves
    68 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @rohanpaul_aiGoogle DeepMind + Harvard + Stanford and many other top labs paper argues that a path to AGI may be visual AI that builds a world model, remembers changes, predicts outcomes, and acts. Most multimodal AI still treats vision as something you feed into a language model. The paper wants vision to do more of the thinking itself. A capable visual system should learn directly from images, video, 3D structure, and interaction. It should understand what exists, what changed, what is hidden, what might happen next, and what it needs to look at before acting. they point to video generation, reconstruction, persistent memory, continual learning, multimodal sensing, and robotics as possible pieces of the same system. they say stop judging visual AI mainly by image Q&A, captions, or realistic video. learn the world's structure from visual experience, keep updating that knowledge, and use it to predict and act.

    3 Sources

    @rohanpaul_aiGoogle DeepMind + Harvard + Stanford and many other top labs paper argues that a path to AGI may be visual AI that builds a world model, remembers changes, predicts outcomes, and acts. Most multimodal AI still treats vision as something you feed into a language model. The paper wants vision to do more of the thinking itself. A capable visual system should learn directly from images, video, 3D structure, and interaction. It should understand what exists, what changed, what is hidden, what might happen next, and what it needs to look at before acting. they point to video generation, reconstruction, persistent memory, continual learning, multimodal sensing, and robotics as possible pieces of the same system. they say stop judging visual AI mainly by image Q&A, captions, or realistic video. learn the world's structure from visual experience, keep updating that knowledge, and use it to predict and act.