• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Gemini Launches Agentic Video Understanding in Flash Models

    Developers can now dynamically process long videos with lower costs and higher accuracy.

    GD
    DH
    SG
    27 Sources, 29d ago, first seen 29d ago

    TLDR

    Google announced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The models can now iteratively scan video timelines, select frame rates, and decide whether to use visual frames, audio, or transcripts. It is available now via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It will come later to the Gemini app and YouTube's Ask YouTube feature. Company posts state the approach uses up to 88 percent fewer tokens, cuts costs by 66 percent, and improves accuracy by about 7 percent over static one-frame-per-second processing.

    Combined views

    1.4M

    27 Sources, first seen 29d ago

    Combined views

    1.4M

    27 Sources, first seen 29d ago

    12.4K likes
    12.4K likes
    480 comments
    5.4K saves
    1.3K reposts

    Sentiment

    Positive91.7%8.3%Negative

    Based on 84 sentiment-bearing replies from 72 accounts across 4 conversations.

    480 comments
    5.4K saves
    1.3K reposts

    Sentiment

    Positive91.7%8.3%Negative

    Based on 84 sentiment-bearing replies from 72 accounts across 4 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    27 Sources

    @GoogleDeepMindWe’re bringing agentic video understanding to our latest Gemini models. They can now analyze videos with better accuracy while using up to 88% fewer tokens. 🧵
    @GoogleWe’re introducing a new capability to our latest Gemini models: agentic video understanding. This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens. See how it works 🧵
    @_philschmidGemini video understanding is now agentic. Gemini can now iteratively navigate video timelines, decide watch what, pick frame rates, or chooses whether it needs speech transcripts, audio, or visual frames to answer your prompt. Result: Long videos get up to 88% fewer tokens and 66% lower costs, with ~7% higher accuracy on benchmarks. How it works: - Receives a lightweight URI reference (Files API or YouTube) and loads content via tool. - Scans speech transcripts to pinpoint relevant moments before fetching visual frames. - Navigates key timestamps and picks its own frame rate (0.1 or 10 FPS). - Pulls audio tracks directly when acoustic cues matter. Set `processing="agentic"`on `video` to enable. Keep `static` (none) for videos under 2 minutes. Available today in the Gemini API and Google AI Studio across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash Lite. http://ai.dev/learn/agentic-video-understanding-with-gemini
    @RohanLikesAIToday, we’re launching Agentic Video Understanding in Gemini across 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. 📉 ~88% fewer tokens 💰 ~66% lower cost 📈 ~7% higher accuracy The result: a new Pareto frontier for video understanding across accuracy and cost. Traditional video understanding processes video statically at a fixed sampling rate. With Agentic Video Understanding, Gemini can actively decide what to watch, at what speed, and which signals to use (frames, audio, or transcripts) by using native video tools. This enables split-second retrieval, rapid-motion understanding, and needle-in-a-haystack search across hours of video. Available now via the Gemini API - and coming soon to billions of users through the Gemini App and YouTube’s “Ask YouTube” experience. Learn more (details, demos, and docs): https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/ PMing this 0→1 has been a career highlight. Taking it from an idea and helping shape it into a frontier-defining capability we’re now shipping broadly has been incredibly rewarding. Huge thanks to @MarioLucic_ , @skprat @ahmetius , @FPavetic, @suhasyogin , @jalayrac, @bcaine, @nbrichtova, @xu_bibo , @tulseedoshi, @davthack, @koraykv and the multimodal team at @GoogleDeepMind.
    @ScobleizerRT @RohanLikesAI: Today, we’re launching Agentic Video Understanding in Gemini across 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. 📉 ~88% fewe…
    @ahmetiusToday, we're launching Agentic Video in Gemini! 🎬🤖 We’re bringing agentic capabilities directly to video in Gemini, enabling models to actively inspect, ground, and reason over temporal visual context. Excited to see what developers and researchers build with this!
    @googleaidevsCan Gemini count the number of claps? 👏 Accurately counting rapid movements is a notoriously tricky task for AI. Because static processing ingests video at a fixed 1 FPS by default, split-second movements like a clap easily get missed entirely or get confused with a snap or click. Watch Gemini 3.7 Flash use the new agentic video understanding capability to accurately identify and count every single clap by automatically adapting the processing speed as needed:
    @osansevieroExcited for all the unlocked use cases. Take a look at the developer guide! https://ai.dev/learn/agentic-video-understanding-with-gemini
    @vamsibatchukprocessing 2-hour long videos used to mean massive token bills. not anymore !!! we are announcing agentic video understanding with Gemini today, that cuts token usage by up to 88% & reduces costs by up to 66% 🔥 another fun launch I got to work with the team. with this capability, you can ask complex questions & create needle-in-a-haystack search apps. for example, users can query multi-hour lectures or how-to guides and the model will dynamically scrub to the exact answer without wasting tokens.
    @OfficialLoganKyou can read more about agentic video in our launch blog: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/ very excited given that Gemini continues to be SOTA in video understanding, pls keep the feedback coming!

    27 Sources

    @GoogleDeepMindWe’re bringing agentic video understanding to our latest Gemini models. They can now analyze videos with better accuracy while using up to 88% fewer tokens. 🧵
    @GoogleWe’re introducing a new capability to our latest Gemini models: agentic video understanding. This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens. See how it works 🧵
    @_philschmidGemini video understanding is now agentic. Gemini can now iteratively navigate video timelines, decide watch what, pick frame rates, or chooses whether it needs speech transcripts, audio, or visual frames to answer your prompt. Result: Long videos get up to 88% fewer tokens and 66% lower costs, with ~7% higher accuracy on benchmarks. How it works: - Receives a lightweight URI reference (Files API or YouTube) and loads content via tool. - Scans speech transcripts to pinpoint relevant moments before fetching visual frames. - Navigates key timestamps and picks its own frame rate (0.1 or 10 FPS). - Pulls audio tracks directly when acoustic cues matter. Set `processing="agentic"`on `video` to enable. Keep `static` (none) for videos under 2 minutes. Available today in the Gemini API and Google AI Studio across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash Lite. http://ai.dev/learn/agentic-video-understanding-with-gemini
    @RohanLikesAIToday, we’re launching Agentic Video Understanding in Gemini across 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. 📉 ~88% fewer tokens 💰 ~66% lower cost 📈 ~7% higher accuracy The result: a new Pareto frontier for video understanding across accuracy and cost. Traditional video understanding processes video statically at a fixed sampling rate. With Agentic Video Understanding, Gemini can actively decide what to watch, at what speed, and which signals to use (frames, audio, or transcripts) by using native video tools. This enables split-second retrieval, rapid-motion understanding, and needle-in-a-haystack search across hours of video. Available now via the Gemini API - and coming soon to billions of users through the Gemini App and YouTube’s “Ask YouTube” experience. Learn more (details, demos, and docs): https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/ PMing this 0→1 has been a career highlight. Taking it from an idea and helping shape it into a frontier-defining capability we’re now shipping broadly has been incredibly rewarding. Huge thanks to @MarioLucic_ , @skprat @ahmetius , @FPavetic, @suhasyogin , @jalayrac, @bcaine, @nbrichtova, @xu_bibo , @tulseedoshi, @davthack, @koraykv and the multimodal team at @GoogleDeepMind.
    @ScobleizerRT @RohanLikesAI: Today, we’re launching Agentic Video Understanding in Gemini across 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. 📉 ~88% fewe…
    @ahmetiusToday, we're launching Agentic Video in Gemini! 🎬🤖 We’re bringing agentic capabilities directly to video in Gemini, enabling models to actively inspect, ground, and reason over temporal visual context. Excited to see what developers and researchers build with this!
    @googleaidevsCan Gemini count the number of claps? 👏 Accurately counting rapid movements is a notoriously tricky task for AI. Because static processing ingests video at a fixed 1 FPS by default, split-second movements like a clap easily get missed entirely or get confused with a snap or click. Watch Gemini 3.7 Flash use the new agentic video understanding capability to accurately identify and count every single clap by automatically adapting the processing speed as needed:
    @osansevieroExcited for all the unlocked use cases. Take a look at the developer guide! https://ai.dev/learn/agentic-video-understanding-with-gemini
    @vamsibatchukprocessing 2-hour long videos used to mean massive token bills. not anymore !!! we are announcing agentic video understanding with Gemini today, that cuts token usage by up to 88% & reduces costs by up to 66% 🔥 another fun launch I got to work with the team. with this capability, you can ask complex questions & create needle-in-a-haystack search apps. for example, users can query multi-hour lectures or how-to guides and the model will dynamically scrub to the exact answer without wasting tokens.
    @OfficialLoganKyou can read more about agentic video in our launch blog: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/ very excited given that Gemini continues to be SOTA in video understanding, pls keep the feedback coming!