• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Announcement

WOVEN training subsets are claimed to collectively improve 22 of 26 downstream benchmarks

A researcher sharing WOVEN says its training recipe selects supervision by reasoning operation, not scene, action or domain similarity.

Mohit BansalMB
Zheyu FanZF
Yue ZhangYZ
4 Sources, 1h ago, first seen 1h ago

TLDR

A researcher sharing WOVEN describes visual transition reasoning as inferring missing information about a state, an action or the resulting state. They say roughly 2,000-example subsets collectively improved 22 of 26 downstream benchmarks by up to 27.3 points. The proposed training recipe selects supervision by the reasoning operation it teaches rather than scene, action or domain similarity.

Combined views

953

4 Sources, first seen 1h ago

18 likes2 comments19 reposts

Combined views

953

4 Sources, first seen 1h ago

18 likes2 comments19 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

4 Sources

Zheyu Fan@ZheyuFan🚨 What makes good world-model training for MLLMs that actually transfers across tasks? 🧶 Excited to share our new work, WOVEN! 🧩 Spatial, physical, temporal, and embodied reasoning all require understanding how the world changes. 🔁 We view world modeling through state transitions p(s′ | s, a), and study visual transition reasoning: inferring what is unobserved in (s, a, s′). 🤔 Can visual transition reasoning be learned as a shared primitive that benefits multiple downstream tasks? 🤔 If so, how should we train it for better downstream gains? Major takeaways: ➡️ Visual transition reasoning is a shared, teachable primitive that transfers widely downstream: ~2K-example subsets collectively improve 22 of 26 downstream benchmarks by up to 27.3 points. 🧪 A data-centric training recipe: to teach this shared primitive for downstream tasks, select supervision by the reasoning operation it teaches, rather than scene/action/domain similarity. Thread🧵👇1h
Mohit Bansal@mohitban47RT @ZheyuFan: 🚨 What makes good world-model training for MLLMs that actually transfers across tasks? 🧶 Excited to share our new work, WOVEN…1h
Yue Zhang@zhan1624Excited to share WOVEN! 🚀 The goal isn't just to teach models what the world looks like. It's to teach them how the world changes. Our key finding: Visual transition reasoning is a shared, teachable primitive that transfers across diverse downstream tasks. And the key to effective transfer? Teach the right reasoning operations, rather than simply matching scenes, actions, or domains.1h
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    4 Sources

    Zheyu Fan@ZheyuFan🚨 What makes good world-model training for MLLMs that actually transfers across tasks? 🧶 Excited to share our new work, WOVEN! 🧩 Spatial, physical, temporal, and embodied reasoning all require understanding how the world changes. 🔁 We view world modeling through state transitions p(s′ | s, a), and study visual transition reasoning: inferring what is unobserved in (s, a, s′). 🤔 Can visual transition reasoning be learned as a shared primitive that benefits multiple downstream tasks? 🤔 If so, how should we train it for better downstream gains? Major takeaways: ➡️ Visual transition reasoning is a shared, teachable primitive that transfers widely downstream: ~2K-example subsets collectively improve 22 of 26 downstream benchmarks by up to 27.3 points. 🧪 A data-centric training recipe: to teach this shared primitive for downstream tasks, select supervision by the reasoning operation it teaches, rather than scene/action/domain similarity. Thread🧵👇1h
    Mohit Bansal@mohitban47RT @ZheyuFan: 🚨 What makes good world-model training for MLLMs that actually transfers across tasks? 🧶 Excited to share our new work, WOVEN…1h
    Yue Zhang@zhan1624Excited to share WOVEN! 🚀 The goal isn't just to teach models what the world looks like. It's to teach them how the world changes. Our key finding: Visual transition reasoning is a shared, teachable primitive that transfers across diverse downstream tasks. And the key to effective transfer? Teach the right reasoning operations, rather than simply matching scenes, actions, or domains.1h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet