• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Could vision-language models 'think with images' without explicit tool calls?

    A post asks whether vision-language models could internalize the effects of calling the right tools in latent space.

    Pasquale MinerviniPM
    1 Source, 1h ago, first seen 1h ago

    TLDR

    The author says tool calls help vision-language models (VLMs) “think with images.” They ask whether VLMs could instead internalize the effects of calling the right tools in latent space, without explicitly calling them during inference. In the October 7 post, the author said they planned to present at COLM the next day at 11 a.m.

    Combined views

    6

    1 Source, first seen 1h ago

    Combined views

    6

    1 Source, first seen 1h ago

    3 reposts
    3 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    Pasquale Minervini@PMinerviniRT @AshutoshAd63687: VLMs benefit from tool calls as they allow for "thinking with images". But can a VLM simply internalise the effects o…1h

    1 Source

    Pasquale Minervini@PMinerviniRT @AshutoshAd63687: VLMs benefit from tool calls as they allow for "thinking with images". But can a VLM simply internalise the effects o…1h