Thaddäus Wiedemer Proposes Visual Prompt Engineering for Video Models
Google DeepMind researcher Thaddäus Wiedemer suggests adapting prompt tweaks to video models on visual reasoning tasks.
TLDR
Thaddäus Wiedemer posted that when an LLM underperforms people tweak its prompt and asked why the same approach should not apply to video models on visual reasoning tasks. He described visual prompt engineering as a method that alters the appearance of the start frame without changing the underlying task and stated that this change improves performance. The post presents the idea as an adaptation of existing prompt engineering practices to a new domain and positions it as a direct suggestion from the author.
Combined views
2.1K
1 Source, first seen 63d ago