• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    a16z Discusses Atlas World Model With World Labs

    a16z shares video of World Labs founders discussing Atlas world model.

    FL
    A1
    SW
    11 Sources, 26d ago, first seen 26d ago

    TLDR

    a16z posted about a 43-minute YouTube video featuring World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall and a16z general partner Martin Casado. The discussion covers Atlas, described as a world model for spatial intelligence built on new view prediction. It contrasts this with LLMs using next token prediction and video models using next frame prediction, noting Atlas as the first to unify pixel generation. The linked video is titled Why World Models Could Change Robotics, 3D, and Creativity.

    Combined views

    1.1M

    11 Sources, first seen 26d ago

    Combined views

    1.1M

    11 Sources, first seen 26d ago

    5.6K likes
    5.6K likes
    177 comments
    2.1K saves
    640 reposts
    177 comments
    2.1K saves
    640 reposts

    Sentiment

    Positive78.6%21.4%Negative

    Summary

    Sentiment

    Positive78.6%21.4%Negative

    Many accounts praised Atlas for its ability to reconstruct 3D scenes directly from photos and recreate effects like Bullet Time using iPhones, while some replies questioned the hype and warned of risks such as surveillance and deepfakes.

    Based on 40 sentiment-bearing replies from 28 accounts across 6 conversations.

    Summary

    Many accounts praised Atlas for its ability to reconstruct 3D scenes directly from photos and recreate effects like Bullet Time using iPhones, while some replies questioned the hype and warned of risks such as surveillance and deepfakes.

    Based on 40 sentiment-bearing replies from 28 accounts across 6 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    11 Sources

    @a16zWorld Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: https://www.youtube.com/watch?v=qn1QDDBnTA0 @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
    @drfeifeiNext view prediction is the key to Atlas, enabling us to unify pixel-level generation and reconstruction. @jcjohnss @BenMildenhall @martin_casado and I had a deeper discussion on some of the most exciting technical innovations of Atlas, our newly released world model for spatial intelligence!
    @sarahdingwangRT @a16z: >be dr. fei-fei li >born in beijing, raised in chengdu >dad moves to new jersey, follow at age 16 >speak almost no english >paren…
    @peteskomorochRT @a16z: World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say LLMs use next-token prediction, but spatial intelligence has its own…

    11 Sources

    @a16zWorld Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: https://www.youtube.com/watch?v=qn1QDDBnTA0 @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
    @drfeifeiNext view prediction is the key to Atlas, enabling us to unify pixel-level generation and reconstruction. @jcjohnss @BenMildenhall @martin_casado and I had a deeper discussion on some of the most exciting technical innovations of Atlas, our newly released world model for spatial intelligence!
    @sarahdingwangRT @a16z: >be dr. fei-fei li >born in beijing, raised in chengdu >dad moves to new jersey, follow at age 16 >speak almost no english >paren…
    @peteskomorochRT @a16z: World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say LLMs use next-token prediction, but spatial intelligence has its own…