• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    ViTeX-Bench debuts as a benchmark for editing text in videos

    Its creators say it includes 387 real-world 720p videos with text-region masks and editing instructions.

    KD
    ZT
    2 Sources, ,

    TLDR

    The creators of ViTeX-Bench say it evaluates video text edits for text correctness, consistency across frames and whether the rest of the scene stays unchanged. The release includes 13 metrics, evaluations of eight baseline models and an open-source reference model. They say their study found that models can preserve a scene yet fail to edit its text, or get the text right while introducing flicker.

    Combined views

    1.5K

    2 Sources, first seen 13h ago

    Combined views

    1.5K

    2 Sources, first seen 13h ago

    24 likes
    13h ago
    first seen 13h ago
    24 likes
    5 saves
    4 reposts

    Sentiment

    Positiveโ€”โ€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    5 saves
    4 reposts

    Sentiment

    Positiveโ€”โ€”Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    โ€”

    Not ranked yet

    Today's Rank

    โ€”

    Not ranked yet

    2 Sources

    @_vztuExcited to share our new work, ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—•๐—ฒ๐—ป๐—ฐ๐—ต: ๐—•๐—ฒ๐—ป๐—ฐ๐—ต๐—บ๐—ฎ๐—ฟ๐—ธ๐—ถ๐—ป๐—ด ๐—›๐—ถ๐—ด๐—ต-๐—™๐—ถ๐—ฑ๐—ฒ๐—น๐—ถ๐˜๐˜† ๐—ฉ๐—ถ๐—ฑ๐—ฒ๐—ผ ๐—ฆ๐—ฐ๐—ฒ๐—ป๐—ฒ ๐—ง๐—ฒ๐˜…๐˜ ๐—˜๐—ฑ๐—ถ๐˜๐—ถ๐—ป๐—ด, accepted to the #NeurIPS 2026 ๐—˜&๐—— ๐—ง๐—ฟ๐—ฎ๐—ฐ๐—ธ! ๐ŸŽ‰ Editing text in videos is surprisingly challenging: a good model needs to generate the correct text, keep it temporally consistent across frames, and preserve everything else in the scene. To study this problem better, we introduce ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—•๐—ฒ๐—ป๐—ฐ๐—ต, a comprehensive benchmark for video scene text editing. Our release includes: ๐Ÿ”น ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐——๐—ฎ๐˜๐—ฎ๐˜€๐—ฒ๐˜ โ€” 387 real-world 720p videos with text-region masks and editing instructions ๐Ÿ”น ๐—” ๐˜๐—ต๐—ฟ๐—ฒ๐—ฒ-๐—ฎ๐˜…๐—ถ๐˜€ ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฝ๐—ฟ๐—ผ๐˜๐—ผ๐—ฐ๐—ผ๐—น covering text correctness, temporal quality, and edit locality with 13 metrics ๐Ÿ”น ๐—˜๐˜…๐˜๐—ฒ๐—ป๐˜€๐—ถ๐˜ƒ๐—ฒ ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ผ๐—ณ ๐Ÿด ๐—ฏ๐—ฎ๐˜€๐—ฒ๐—น๐—ถ๐—ป๐—ฒ๐˜€ across four video-editing paradigms ๐Ÿ”น ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—˜๐—ฑ๐—ถ๐˜-๐Ÿญ๐Ÿฐ๐—•, an open-source reference model with motion-aligned glyph-video conditioning ๐Ÿ”น Open-source ๐—ฑ๐—ฎ๐˜๐—ฎ, ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฐ๐—ผ๐—ฑ๐—ฒ, ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐˜„๐—ฒ๐—ถ๐—ด๐—ต๐˜๐˜€, ๐—ฎ๐—ป๐—ฑ ๐—น๐—ฒ๐—ฎ๐—ฑ๐—ฒ๐—ฟ๐—ฏ๐—ผ๐—ฎ๐—ฟ๐—ฑ One takeaway from our study: no single metric tells the whole story. Models that preserve the scene well may fail to actually edit the text, while models that generate the right text frame-by-frame can still suffer from severe temporal flickering. ViTeX-Bench is designed to make these trade-offs explicit and reproducible. We hope ViTeX-Bench can provide a useful foundation for future research on precise and temporally consistent video editing. ๐Ÿ”— Project page: http://vitex-bench.github.io ๐Ÿ“œ Paper link: http://arxiv.org/abs/2609.40356 #NeurIPS2026 #ComputerVision #GenerativeAI #VideoEditing #VideoGeneration #MultimodalAI #MachineLearning #AIResearch
    @CSProfKGDRT @_vztu: Excited to share our new work, ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—•๐—ฒ๐—ป๐—ฐ๐—ต: ๐—•๐—ฒ๐—ป๐—ฐ๐—ต๐—บ๐—ฎ๐—ฟ๐—ธ๐—ถ๐—ป๐—ด ๐—›๐—ถ๐—ด๐—ต-๐—™๐—ถ๐—ฑ๐—ฒ๐—น๐—ถ๐˜๐˜† ๐—ฉ๐—ถ๐—ฑ๐—ฒ๐—ผ ๐—ฆ๐—ฐ๐—ฒ๐—ป๐—ฒ ๐—ง๐—ฒ๐˜…๐˜ ๐—˜๐—ฑ๐—ถ๐˜๐—ถ๐—ป๐—ด, accepted to the #NeurIPS 2026 ๐—˜โ€ฆ

    2 Sources

    @_vztuExcited to share our new work, ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—•๐—ฒ๐—ป๐—ฐ๐—ต: ๐—•๐—ฒ๐—ป๐—ฐ๐—ต๐—บ๐—ฎ๐—ฟ๐—ธ๐—ถ๐—ป๐—ด ๐—›๐—ถ๐—ด๐—ต-๐—™๐—ถ๐—ฑ๐—ฒ๐—น๐—ถ๐˜๐˜† ๐—ฉ๐—ถ๐—ฑ๐—ฒ๐—ผ ๐—ฆ๐—ฐ๐—ฒ๐—ป๐—ฒ ๐—ง๐—ฒ๐˜…๐˜ ๐—˜๐—ฑ๐—ถ๐˜๐—ถ๐—ป๐—ด, accepted to the #NeurIPS 2026 ๐—˜&๐—— ๐—ง๐—ฟ๐—ฎ๐—ฐ๐—ธ! ๐ŸŽ‰ Editing text in videos is surprisingly challenging: a good model needs to generate the correct text, keep it temporally consistent across frames, and preserve everything else in the scene. To study this problem better, we introduce ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—•๐—ฒ๐—ป๐—ฐ๐—ต, a comprehensive benchmark for video scene text editing. Our release includes: ๐Ÿ”น ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐——๐—ฎ๐˜๐—ฎ๐˜€๐—ฒ๐˜ โ€” 387 real-world 720p videos with text-region masks and editing instructions ๐Ÿ”น ๐—” ๐˜๐—ต๐—ฟ๐—ฒ๐—ฒ-๐—ฎ๐˜…๐—ถ๐˜€ ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฝ๐—ฟ๐—ผ๐˜๐—ผ๐—ฐ๐—ผ๐—น covering text correctness, temporal quality, and edit locality with 13 metrics ๐Ÿ”น ๐—˜๐˜…๐˜๐—ฒ๐—ป๐˜€๐—ถ๐˜ƒ๐—ฒ ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ผ๐—ณ ๐Ÿด ๐—ฏ๐—ฎ๐˜€๐—ฒ๐—น๐—ถ๐—ป๐—ฒ๐˜€ across four video-editing paradigms ๐Ÿ”น ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—˜๐—ฑ๐—ถ๐˜-๐Ÿญ๐Ÿฐ๐—•, an open-source reference model with motion-aligned glyph-video conditioning ๐Ÿ”น Open-source ๐—ฑ๐—ฎ๐˜๐—ฎ, ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ฐ๐—ผ๐—ฑ๐—ฒ, ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐˜„๐—ฒ๐—ถ๐—ด๐—ต๐˜๐˜€, ๐—ฎ๐—ป๐—ฑ ๐—น๐—ฒ๐—ฎ๐—ฑ๐—ฒ๐—ฟ๐—ฏ๐—ผ๐—ฎ๐—ฟ๐—ฑ One takeaway from our study: no single metric tells the whole story. Models that preserve the scene well may fail to actually edit the text, while models that generate the right text frame-by-frame can still suffer from severe temporal flickering. ViTeX-Bench is designed to make these trade-offs explicit and reproducible. We hope ViTeX-Bench can provide a useful foundation for future research on precise and temporally consistent video editing. ๐Ÿ”— Project page: http://vitex-bench.github.io ๐Ÿ“œ Paper link: http://arxiv.org/abs/2609.40356 #NeurIPS2026 #ComputerVision #GenerativeAI #VideoEditing #VideoGeneration #MultimodalAI #MachineLearning #AIResearch
    @CSProfKGDRT @_vztu: Excited to share our new work, ๐—ฉ๐—ถ๐—ง๐—ฒ๐—ซ-๐—•๐—ฒ๐—ป๐—ฐ๐—ต: ๐—•๐—ฒ๐—ป๐—ฐ๐—ต๐—บ๐—ฎ๐—ฟ๐—ธ๐—ถ๐—ป๐—ด ๐—›๐—ถ๐—ด๐—ต-๐—™๐—ถ๐—ฑ๐—ฒ๐—น๐—ถ๐˜๐˜† ๐—ฉ๐—ถ๐—ฑ๐—ฒ๐—ผ ๐—ฆ๐—ฐ๐—ฒ๐—ป๐—ฒ ๐—ง๐—ฒ๐˜…๐˜ ๐—˜๐—ฑ๐—ถ๐˜๐—ถ๐—ป๐—ด, accepted to the #NeurIPS 2026 ๐—˜โ€ฆ