Announcement
WOVEN training subsets are claimed to collectively improve 22 of 26 downstream benchmarks
A researcher sharing WOVEN says its training recipe selects supervision by reasoning operation, not scene, action or domain similarity.
TLDR
A researcher sharing WOVEN describes visual transition reasoning as inferring missing information about a state, an action or the resulting state. They say roughly 2,000-example subsets collectively improved 22 of 26 downstream benchmarks by up to 27.3 points. The proposed training recipe selects supervision by the reasoning operation it teaches rather than scene, action or domain similarity.
Combined views
953
4 Sources, first seen ago
18 likes2 comments19 reposts
