VideoTreeSearch Organizes Videos As Trees For Self-Correcting QA Agents
Reactions from ranked influencers
2 postsExcited to share our new work, “Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA.” Grounded LVQA requires a model to answer a question about a long video and localize the short interval supporting its answer. We introduce VideoTreeSearch (VTS), which organizes the video as an adaptive temporal tree and lets an agent search, backtrack, and revise its search path. VTS brings strong gains on three grounded LVQA benchmarks and also transfers to general long-video QA. (1/9)
Our proposed VideoTreeSearch achieves state-of-the-art results across all grounded long-video QA benchmarks while also being more efficient than prior agentic video methods. We also show that we can achieve strong temporal grounding results even if we don't use any manually labeled long video data. Check out Ce's thread below for more details. Work with @cezhhh @YuluPan_00 @ZiyangW00 @mohitban47
Excited to share our new work, “Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA.” Grounded LVQA requires a model to answer a question about a long video and localize the short interval supporting its answer. We introduce VideoTreeSearch (VTS), which organizes the video as an adaptive temporal tree and lets an agent search, backtrack, and revise its search path. VTS brings strong gains on three grounded LVQA benchmarks and also transfers to general long-video QA. (1/9)
Combined views
13
2 posts, first seen 2h ago