• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    World Mechanics is being co-founded to develop interpretable models of the physical world

    A co-founder says the lab wants to make interpretability an objective during training, rather than analyze models only afterward.

    IR
    SJ
    CV
    5 Sources, ,

    TLDR

    A co-founder announced World Mechanics, a frontier lab they are starting with another researcher to work on interpretable foundation models for the physical world. They argue that interventions in simulations or on real machines could help researchers understand what the models learn, and say the lab wants to do as much work openly as possible.

    Combined views

    10.9K

    5 Sources, first seen 3h ago

    Combined views

    10.9K

    5 Sources, first seen 3h ago

    130 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3h ago
    first seen 3h ago
    130 likes
    13 comments
    42 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    13 comments
    42 saves
    34 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 Sources

    @cvenhoff00I’m excited to share that I am co-founding World Mechanics, a frontier lab working on interpretable foundation models for the physical world, together with @soniajoseph_. A big reason I’m excited about physical AI is that it gives us a rare opportunity to rethink interpretability from the ground up. For LLMs, much of interpretability is necessarily post hoc, where we take a model that has already been trained and try to reverse-engineer its internal representations, often without clear ground truth for what it should have learned. For physical AI, we’re still early enough to make interpretability a first-order objective of training itself. Physical systems also give us unusually useful data-generating processes for interpretability. We can intervene on systems in simulation or on real machines, observe how they respond, and make use of known physical structure. This gives us a path toward representations that faithfully abstract the underlying system and remain controllable under intervention. At the same time, the interpretability problem itself changes. Video and sensor data are not directly legible in the way language is, and models that reason about the physical world learn very different representations from those we see in LLMs. Many interpretability methods developed for language models already do not transfer directly, leaving a lot of fundamental work to do. This matters because safe and generalizable behavior of physical AI models depends on them learning faithful abstractions of the systems they act on, including the right semantics and causal structure. Current black-box evaluations give us only indirect evidence that these abstractions are correct, since they cannot cover every scenario or explain why a model fails in a particular one. Interpretability-based white-box evaluations, on the other hand, would allow us to understand the model’s representations directly, explain failure modes, and use those insights to improve the reliability and safety of physical AI models. Because the field is still early, I also care a lot about how the research community around it develops. We want to do as much of our work in the open as possible and collaborate closely with researchers in academia and elsewhere. If you’re working on related problems, or this sounds like something you’d want to help build, reach out!
    @soniajoseph_@cvenhoff00 @MATSprogram @AIatMeta Fun fact: we were introduced by @NeelNanda5 two years ago in the context of interpretablity VLM research!
    @philiptorrAwesome news from my student Constantin...
    @irinarishRT @soniajoseph_: I’ve collaborated with @cvenhoff00 for 2 years across institutions, from VLM failure traces at @MATSprogram to interpreta…

    5 Sources

    @cvenhoff00I’m excited to share that I am co-founding World Mechanics, a frontier lab working on interpretable foundation models for the physical world, together with @soniajoseph_. A big reason I’m excited about physical AI is that it gives us a rare opportunity to rethink interpretability from the ground up. For LLMs, much of interpretability is necessarily post hoc, where we take a model that has already been trained and try to reverse-engineer its internal representations, often without clear ground truth for what it should have learned. For physical AI, we’re still early enough to make interpretability a first-order objective of training itself. Physical systems also give us unusually useful data-generating processes for interpretability. We can intervene on systems in simulation or on real machines, observe how they respond, and make use of known physical structure. This gives us a path toward representations that faithfully abstract the underlying system and remain controllable under intervention. At the same time, the interpretability problem itself changes. Video and sensor data are not directly legible in the way language is, and models that reason about the physical world learn very different representations from those we see in LLMs. Many interpretability methods developed for language models already do not transfer directly, leaving a lot of fundamental work to do. This matters because safe and generalizable behavior of physical AI models depends on them learning faithful abstractions of the systems they act on, including the right semantics and causal structure. Current black-box evaluations give us only indirect evidence that these abstractions are correct, since they cannot cover every scenario or explain why a model fails in a particular one. Interpretability-based white-box evaluations, on the other hand, would allow us to understand the model’s representations directly, explain failure modes, and use those insights to improve the reliability and safety of physical AI models. Because the field is still early, I also care a lot about how the research community around it develops. We want to do as much of our work in the open as possible and collaborate closely with researchers in academia and elsewhere. If you’re working on related problems, or this sounds like something you’d want to help build, reach out!
    @soniajoseph_@cvenhoff00 @MATSprogram @AIatMeta Fun fact: we were introduced by @NeelNanda5 two years ago in the context of interpretablity VLM research!
    @philiptorrAwesome news from my student Constantin...
    @irinarishRT @soniajoseph_: I’ve collaborated with @cvenhoff00 for 2 years across institutions, from VLM failure traces at @MATSprogram to interpreta…