Gemini Launches Agentic Video Understanding in Flash Models
Developers can now dynamically process long videos with lower costs and higher accuracy.
TLDR
Google announced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The models can now iteratively scan video timelines, select frame rates, and decide whether to use visual frames, audio, or transcripts. It is available now via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It will come later to the Gemini app and YouTube's Ask YouTube feature. Company posts state the approach uses up to 88 percent fewer tokens, cuts costs by 66 percent, and improves accuracy by about 7 percent over static one-frame-per-second processing.
Combined views
1.4M
27 Sources, first seen 29d ago