• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Counterfactual Debugging Scales for World Model Agents

    Columbia PhD candidate presents method using causal attribution to find sim2real gaps.

    MD
    ML
    3 Sources, 28d ago, first seen 28d ago

    TLDR

    Mingxuan Li posted that agents trained in world models often fail after deployment and that identifying the true sim2real gap remains difficult. The proposed Counterfactual Debugging approach uses causal attribution to locate the cause. Michael Dennis, a DeepMind researcher, retweeted the thread and stated that world models create new safety options such as simulated counterfactuals for diagnosing failures. He added that the presented work applies divide and conquer to extend the technique.

    Combined views

    5K

    3 Sources, first seen 28d ago

    Combined views

    5K

    3 Sources, first seen 28d ago

    35 likes
    35 likes
    1 comments
    23 saves
    7 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    23 saves
    7 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 Sources

    @Mingxuan0422World model trained agents often fail in deployment. Identifying the real sim2real gap is hard. Our solution, Counterfactual Debugging, pinpoints the cause via causal attribution at 1M steps. 🧵(1/8)
    @MichaelD1729RT @Mingxuan0422: World model trained agents often fail in deployment. Identifying the real sim2real gap is hard. Our solution, Counterfact…

    3 Sources

    @Mingxuan0422World model trained agents often fail in deployment. Identifying the real sim2real gap is hard. Our solution, Counterfactual Debugging, pinpoints the cause via causal attribution at 1M steps. 🧵(1/8)
    @MichaelD1729RT @Mingxuan0422: World model trained agents often fail in deployment. Identifying the real sim2real gap is hard. Our solution, Counterfact…