• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    The case for language models at the core of physical AI

    A researcher describes work across Gemini, Omni and generative media to test ideas about physical AI, arguing that language models remain indispensable.

    Shane GuSG
    David PfauDP
    Angjoo KanazawaAK
    16 Sources, ,

    TLDR

    One post describes unease among academics over what it calls a leap in robotics and world-model benchmark performance by Astra, Fable and Muse. A researcher quoting that post argues that digital, symbolic artificial general intelligence must precede—and would accelerate—physical AI. They describe language models as indispensable to that development and say frontier labs, alongside the broader robotics community, are equipped to drive the next transition.

    Combined views

    853.5K

    16 Sources, first seen 22d ago

    Combined views

    853.5K

    16 Sources, first seen 22d ago

    4K likes
    22d ago
    first seen 22d ago
    4K likes
    362 comments
    1.7K saves
    405 reposts
    362 comments
    1.7K saves
    405 reposts

    Sentiment

    Positive47.6%52.4%Negative

    Summary

    Positive accounts value extended PhD research time and AI tools for industrial use, while negative replies question academic relevance, cite plagiarism risks in VLMs like Astra, and criticize publish-or-perish incentives.

    Based on 22 sentiment-bearing replies from 21 accounts across 3 conversations.

    Sentiment

    Positive47.6%52.4%Negative

    Summary

    Positive accounts value extended PhD research time and AI tools for industrial use, while negative replies question academic relevance, cite plagiarism risks in VLMs like Astra, and criticize publish-or-perish incentives.

    Based on 22 sentiment-bearing replies from 21 accounts across 3 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    16 Sources

    Letian Wang@Letian_Wang_6Earlier, quite a few people who attended CVPR told me they came away pretty disappointed — including some extremely senior people in the field. At the time, I didn’t fully understand why. But after I attended ECCV, I finally did. There was a surprising amount of frustration at the conference, especially since it happened right after the GPT-6 release. People were openly wondering what traditional CV research is going to look like from here, and how much of it will eventually just be absorbed by large foundation models. Actually, when I gave talks about GenCeption to a few academic groups, I had already started to see that kind of frustration. GenCeption essentially leverages a large-scale pretrained diffusion model to build a single unified model that can match — and often outperform — specialized SOTA models across a surprisingly broad range of vision tasks, including DepthAnything3, D4RT, VGGT Omega, SAM3, Genmo, Lotus-2, and others. (The code is now open-sourced! https://genception.github.io/) Quite a few PhD students told me the work made them question the direction of their own research. Part of me was happy, because it meant GenCeption had genuinely changed how some people thought about the problem. But that feeling was quickly outweighed by how much I felt bad and empathized with the junior PhD students. Some of them seemed genuinely unsure where to go next, especially when years of work on specialized architectures or training recipes could potentially be absorbed by a sufficiently capable general pretrained model. GPT-6 seems to be just creating that same feeling on a much broader scale. It reminds me a lot of what happened with GPT-1 and BERT — tons of NLP PhDs' years of work and entire research directions perish overnight. There were even some pretty spicy takes floating around: “CV conferences are dead,” “academia is becoming disconnected from where some of the most important progress is actually happening,” “with increasingly noisy review systems and declining review quality, rejection sometimes is almost a signal of truly fundamental work,” and “conferences increasingly reward work delicately engineered around established topics or methodologies, rather than work truly at the frontier.” Maybe it's too harsh and too absolute, but the frustration behind them was definitely real. Conferences used to feel so fruitful, but this time it was hard to get genuinely excited. A lot of papers seemed to solve problems that no longer felt that important, or make progress that felt increasingly incremental. In the era of paper inflation, most people are busy trying to survive the research game, rather than thinking about how to go big after one or two years of diving deep into something. Most valuable skills such as critical thinking and taste never get trained well enough, from the beginning, all the way to the end. On the other hand, most people simply don't have enough context or a big enough picture about what problem actually matters. They get stuck optimizing for local problems that their immediate environment tells them are important. I suffered from this for a long time too, but luckily I gradually started to see a much bigger picture. My advice to junior PhD students is: when you start a project, try to collect as much context as possible before you even begin — by all means. One way is to explore the boundaries of existing models as aggressively as you can. Play with the SOTA models extensively. Keep throwing harder and harder problems at them until they fundamentally break. Understand not just what they can do today, but where their capabilities end and why. Another way is to put yourself wherever the strongest signals are. That could be a visionary academic lab, a team at an industry frontier lab, a genuinely insightful seminar, or simply a senior student or researcher who has seen a much broader landscape than you have. Expand your local context as much as possible. In an era where things are moving this fast, the information gap alone can be large enough to make years of carefully built work suddenly meaningless. Looking up and making sure you're walking in the right direction matters much more than keeping your head down and perfecting every step along the way — though, of course, being able to do both is the best :) We stand on the shoulders of giants not just to go further, but to see further — to understand which problems will actually matter in the future. If a problem is likely to be easily swallowed by large models once they have sufficient data, attention, or business interest, either avoid spending years engineering a specialized solution for it, or contribute the data and benchmarks that will help move the frontier forward faster. The problems that deserve the deepest thinking and genuinely new methods are the ones where existing models fundamentally cannot do the thing you need them to do. What I actually worry about is not that conferences will suddenly “die,” but that if they start becoming genuinely irrelevant to where the frontier is moving, they will simply fade away. And maybe one early sign of that is that a lot of people I used to want to meet aren’t really coming to conferences anymore — they’re busy with the big things happening at frontier labs, or, well, busy making money. It reminds me a little of what happened to control theory as the field matured. As many of the core problems were gradually solved, parts of the community started inventing increasingly artificial use cases, making increasingly unrealistic assumptions, and playing increasingly elaborate mathematical games within them. That possibility genuinely scares me — not because the work becomes unnecessarily sophisticated, but because it can become disconnected from problems that actually matter. I still believe that conferences will remain places people genuinely want to come to: to exchange ideas, reconnect with old friends, meet new people, and simply have fun together. We just need to keep doing the right things to make that happen :)22d
    Shane Gu@shaneguMLIt’s great seeing so many frontier lab folks, VCs, and startups getting excited about robotics. Always happy to chat and offer alternative perspectives. Before focusing on ChatGPT and Gemini over the past 4 years, I: - Was a founding intern at Google Brain Robotics, first-authoring one of its first 3 papers: https://ieeexplore.ieee.org/abstract/document/7989385 - Co-led a robot dexterity moonshot in 2019 (I avoided non-dexterous tasks, feeling they were ultimately VLM problems) - Won Best Paper at CoRL 2019 https://2019.corl.org/ (the 1st author is now a founding engineer at Generalist) - Angel invested in Sunday / Eka - Was an avid MuJoCo hacker (and dabbled in Brax) My DMs are open!22d
    Alexander Doria@Dorialexanderif you wondered why the whole discourse "europe is betting on the new ai frontiers" suddenly dropped dead21d
    Lerrel Pinto@LerrelPintoRT @shaneguML: I felt a similar despair 4-5 years ago as I was still partially working on robotics. After co-authoring "LLMs are Zero-Shot…21d
    Angjoo Kanazawa@akanazawaI see a lot of doom & gloom around 3D/CV lately. Regardless of the credit assignment, it’s exciting to see the problems we’ve been working on move forward! It does suck that the labs are closed. But don't let that defeat your spirit, ask the right questions and share your work!21d
    Kosta Derpanis@CSProfKGDRT @Letian_Wang_6: Earlier, quite a few people who attended CVPR told me they came away pretty disappointed — including some extremely seni…21d
    Georgia Gkioxari@georgiagkioxari1. Problems that can be solved by scale because data exists at scale will be solved by large-scale approaches like Astra. 2. Producing data at scale is the challenge. 3D/4D was not available at scale; it only became so because researchers in the past 10 years unlocked it (generative 3D, GS, etc) -- same thing for quasi-static manipulation. 3. There are **many** unsolved problems that we haven't gotten to scale yet; that's the next frontier of problems to go after.21d
    Ravid Shwartz Ziv@ziv_ravidMy 2 cents: I’ve met very few people who regret spending a few more years doing research, and many who regret leaving research too early. The boundary between research and product is much thinner today than it used to be, especially in AI. You can move from research into products, startups, or industry later. Going in the other direction is often harder. So if you’re genuinely unsure, I’d stay close to research a little longer. It keeps more doors open, and in a field changing this quickly, optionality is valuable.21d
    Mariya I. Vasileva@mariyaivasilevaRT @shaneguML: I felt a similar despair 4-5 years ago as I was still partially working on robotics. After co-authoring "LLMs are Zero-Shot…20d
    Kostas Daniilidis@KostasPenn"These systems build on the tools we’ve built and what we’ve learned." @akanazawa This is the best response to the relevance discussion initiated by @Letian_Wang_6 and @Michael_J_Black. In the agentic age of robotics you will be relevant if your method will be picked up by an agent. I am thrilled that agents explain what they do and what tools they picked, and it is becoming even more crucial to provide sound methods and as Terrance Tao said to understand how the agent reasons. Of course, we will not compete with the agents in performance. We will compete on what methods or model the agent picked to achieve their goals and how can we provide guarantees or bounds.20d

    16 Sources

    Letian Wang@Letian_Wang_6Earlier, quite a few people who attended CVPR told me they came away pretty disappointed — including some extremely senior people in the field. At the time, I didn’t fully understand why. But after I attended ECCV, I finally did. There was a surprising amount of frustration at the conference, especially since it happened right after the GPT-6 release. People were openly wondering what traditional CV research is going to look like from here, and how much of it will eventually just be absorbed by large foundation models. Actually, when I gave talks about GenCeption to a few academic groups, I had already started to see that kind of frustration. GenCeption essentially leverages a large-scale pretrained diffusion model to build a single unified model that can match — and often outperform — specialized SOTA models across a surprisingly broad range of vision tasks, including DepthAnything3, D4RT, VGGT Omega, SAM3, Genmo, Lotus-2, and others. (The code is now open-sourced! https://genception.github.io/) Quite a few PhD students told me the work made them question the direction of their own research. Part of me was happy, because it meant GenCeption had genuinely changed how some people thought about the problem. But that feeling was quickly outweighed by how much I felt bad and empathized with the junior PhD students. Some of them seemed genuinely unsure where to go next, especially when years of work on specialized architectures or training recipes could potentially be absorbed by a sufficiently capable general pretrained model. GPT-6 seems to be just creating that same feeling on a much broader scale. It reminds me a lot of what happened with GPT-1 and BERT — tons of NLP PhDs' years of work and entire research directions perish overnight. There were even some pretty spicy takes floating around: “CV conferences are dead,” “academia is becoming disconnected from where some of the most important progress is actually happening,” “with increasingly noisy review systems and declining review quality, rejection sometimes is almost a signal of truly fundamental work,” and “conferences increasingly reward work delicately engineered around established topics or methodologies, rather than work truly at the frontier.” Maybe it's too harsh and too absolute, but the frustration behind them was definitely real. Conferences used to feel so fruitful, but this time it was hard to get genuinely excited. A lot of papers seemed to solve problems that no longer felt that important, or make progress that felt increasingly incremental. In the era of paper inflation, most people are busy trying to survive the research game, rather than thinking about how to go big after one or two years of diving deep into something. Most valuable skills such as critical thinking and taste never get trained well enough, from the beginning, all the way to the end. On the other hand, most people simply don't have enough context or a big enough picture about what problem actually matters. They get stuck optimizing for local problems that their immediate environment tells them are important. I suffered from this for a long time too, but luckily I gradually started to see a much bigger picture. My advice to junior PhD students is: when you start a project, try to collect as much context as possible before you even begin — by all means. One way is to explore the boundaries of existing models as aggressively as you can. Play with the SOTA models extensively. Keep throwing harder and harder problems at them until they fundamentally break. Understand not just what they can do today, but where their capabilities end and why. Another way is to put yourself wherever the strongest signals are. That could be a visionary academic lab, a team at an industry frontier lab, a genuinely insightful seminar, or simply a senior student or researcher who has seen a much broader landscape than you have. Expand your local context as much as possible. In an era where things are moving this fast, the information gap alone can be large enough to make years of carefully built work suddenly meaningless. Looking up and making sure you're walking in the right direction matters much more than keeping your head down and perfecting every step along the way — though, of course, being able to do both is the best :) We stand on the shoulders of giants not just to go further, but to see further — to understand which problems will actually matter in the future. If a problem is likely to be easily swallowed by large models once they have sufficient data, attention, or business interest, either avoid spending years engineering a specialized solution for it, or contribute the data and benchmarks that will help move the frontier forward faster. The problems that deserve the deepest thinking and genuinely new methods are the ones where existing models fundamentally cannot do the thing you need them to do. What I actually worry about is not that conferences will suddenly “die,” but that if they start becoming genuinely irrelevant to where the frontier is moving, they will simply fade away. And maybe one early sign of that is that a lot of people I used to want to meet aren’t really coming to conferences anymore — they’re busy with the big things happening at frontier labs, or, well, busy making money. It reminds me a little of what happened to control theory as the field matured. As many of the core problems were gradually solved, parts of the community started inventing increasingly artificial use cases, making increasingly unrealistic assumptions, and playing increasingly elaborate mathematical games within them. That possibility genuinely scares me — not because the work becomes unnecessarily sophisticated, but because it can become disconnected from problems that actually matter. I still believe that conferences will remain places people genuinely want to come to: to exchange ideas, reconnect with old friends, meet new people, and simply have fun together. We just need to keep doing the right things to make that happen :)22d
    Shane Gu@shaneguMLIt’s great seeing so many frontier lab folks, VCs, and startups getting excited about robotics. Always happy to chat and offer alternative perspectives. Before focusing on ChatGPT and Gemini over the past 4 years, I: - Was a founding intern at Google Brain Robotics, first-authoring one of its first 3 papers: https://ieeexplore.ieee.org/abstract/document/7989385 - Co-led a robot dexterity moonshot in 2019 (I avoided non-dexterous tasks, feeling they were ultimately VLM problems) - Won Best Paper at CoRL 2019 https://2019.corl.org/ (the 1st author is now a founding engineer at Generalist) - Angel invested in Sunday / Eka - Was an avid MuJoCo hacker (and dabbled in Brax) My DMs are open!22d
    Alexander Doria@Dorialexanderif you wondered why the whole discourse "europe is betting on the new ai frontiers" suddenly dropped dead21d
    Lerrel Pinto@LerrelPintoRT @shaneguML: I felt a similar despair 4-5 years ago as I was still partially working on robotics. After co-authoring "LLMs are Zero-Shot…21d
    Angjoo Kanazawa@akanazawaI see a lot of doom & gloom around 3D/CV lately. Regardless of the credit assignment, it’s exciting to see the problems we’ve been working on move forward! It does suck that the labs are closed. But don't let that defeat your spirit, ask the right questions and share your work!21d
    Kosta Derpanis@CSProfKGDRT @Letian_Wang_6: Earlier, quite a few people who attended CVPR told me they came away pretty disappointed — including some extremely seni…21d
    Georgia Gkioxari@georgiagkioxari1. Problems that can be solved by scale because data exists at scale will be solved by large-scale approaches like Astra. 2. Producing data at scale is the challenge. 3D/4D was not available at scale; it only became so because researchers in the past 10 years unlocked it (generative 3D, GS, etc) -- same thing for quasi-static manipulation. 3. There are **many** unsolved problems that we haven't gotten to scale yet; that's the next frontier of problems to go after.21d
    Ravid Shwartz Ziv@ziv_ravidMy 2 cents: I’ve met very few people who regret spending a few more years doing research, and many who regret leaving research too early. The boundary between research and product is much thinner today than it used to be, especially in AI. You can move from research into products, startups, or industry later. Going in the other direction is often harder. So if you’re genuinely unsure, I’d stay close to research a little longer. It keeps more doors open, and in a field changing this quickly, optionality is valuable.21d
    Mariya I. Vasileva@mariyaivasilevaRT @shaneguML: I felt a similar despair 4-5 years ago as I was still partially working on robotics. After co-authoring "LLMs are Zero-Shot…20d
    Kostas Daniilidis@KostasPenn"These systems build on the tools we’ve built and what we’ve learned." @akanazawa This is the best response to the relevance discussion initiated by @Letian_Wang_6 and @Michael_J_Black. In the agentic age of robotics you will be relevant if your method will be picked up by an agent. I am thrilled that agents explain what they do and what tools they picked, and it is becoming even more crucial to provide sound methods and as Terrance Tao said to understand how the agent reasons. Of course, we will not compete with the agents in performance. We will compete on what methods or model the agent picked to achieve their goals and how can we provide guarantees or bounds.20d