Counterfactual Debugging Scales for World Model Agents
Columbia PhD candidate presents method using causal attribution to find sim2real gaps.
TLDR
Mingxuan Li posted that agents trained in world models often fail after deployment and that identifying the true sim2real gap remains difficult. The proposed Counterfactual Debugging approach uses causal attribution to locate the cause. Michael Dennis, a DeepMind researcher, retweeted the thread and stated that world models create new safety options such as simulated counterfactuals for diagnosing failures. He added that the presented work applies divide and conquer to extend the technique.
Combined views
5K
3 Sources, first seen 28d ago