Could vision-language models 'think with images' without explicit tool calls?
A post asks whether vision-language models could internalize the effects of calling the right tools in latent space.
TLDR
The author says tool calls help vision-language models (VLMs) “think with images.” They ask whether VLMs could instead internalize the effects of calling the right tools in latent space, without explicitly calling them during inference. In the October 7 post, the author said they planned to present at COLM the next day at 11 a.m.
Combined views
6
1 Source, first seen ago
