• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    PhD Student Introduces VRRL for VLM Visual Self-Reflection

    Tweet from NYU researcher Fangcong Yin presents reinforcement learning method for vision-language models.

    GD
    FY
    2 Sources, 27d ago, first seen 27d ago

    TLDR

    Fangcong Yin announced VRRL, a reinforcement learning approach designed to help vision-language models self-reflect using visual feedback. The post highlights that these models often fail to correct errors despite visual indications of mistakes. It reports that VRRL delivers improved out-of-distribution accuracy in visually grounded reasoning tasks relative to conventional reinforcement learning and supervised fine-tuning techniques. The announcement includes an image with the partial title Visually Grounded Self-Reflection for Vision-Language Models via Reinforceme.

    Combined views

    8.3K

    2 Sources, first seen 27d ago

    Combined views

    8.3K

    2 Sources, first seen 27d ago

    83 likes
    83 likes
    3 comments
    37 saves
    22 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    37 saves
    22 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @fangcong_y10593🤯 Vision-language models fail to correct their mistakes, even when a visual environment shows what went wrong! 🕶️Introducing VRRL, an RL method that teaches VLMs to self-reflect based on visual feedback. Stronger OOD acc. in visually grounded reasoning than RL & SFT methods.
    @gregd_nlpNew work from Liyan and Fangcong on visual self-correction: given an image, a VLM uses a tool to annotate the image, then looks at what it did, and iteratively corrects and updates. VLMs struggle to do this well, but with the right RL paradigm + reward, they can do better!

    2 Sources

    @fangcong_y10593🤯 Vision-language models fail to correct their mistakes, even when a visual environment shows what went wrong! 🕶️Introducing VRRL, an RL method that teaches VLMs to self-reflect based on visual feedback. Stronger OOD acc. in visually grounded reasoning than RL & SFT methods.
    @gregd_nlpNew work from Liyan and Fangcong on visual self-correction: given an image, a VLM uses a tool to annotate the image, then looks at what it did, and iteratively corrects and updates. VLMs struggle to do this well, but with the right RL paradigm + reward, they can do better!