• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Thaddäus Wiedemer Proposes Visual Prompt Engineering for Video Models

    Google DeepMind researcher Thaddäus Wiedemer suggests adapting prompt tweaks to video models on visual reasoning tasks.

    TW
    1 Source, 63d ago, first seen 63d ago

    TLDR

    Thaddäus Wiedemer posted that when an LLM underperforms people tweak its prompt and asked why the same approach should not apply to video models on visual reasoning tasks. He described visual prompt engineering as a method that alters the appearance of the start frame without changing the underlying task and stated that this change improves performance. The post presents the idea as an adaptation of existing prompt engineering practices to a new domain and positions it as a direct suggestion from the author.

    Combined views

    2.1K

    1 Source, first seen 63d ago

    Combined views

    2.1K

    1 Source, first seen 63d ago

    15 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    15 likes
    2 comments
    6 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 comments
    6 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @thwiedemerWhen an LLM is underperforming, we tweak the prompt. Why not do the same for video models solving visual reasoning tasks? Visual prompt engineering changes the start frame appearance but not the task, improving performance. 📄 https://arxiv.org/pdf/2607.25537 🌐 https://visual-prompt-engineering.github.io/

    1 Source

    @thwiedemerWhen an LLM is underperforming, we tweak the prompt. Why not do the same for video models solving visual reasoning tasks? Visual prompt engineering changes the start frame appearance but not the task, improving performance. 📄 https://arxiv.org/pdf/2607.25537 🌐 https://visual-prompt-engineering.github.io/