Gemini Launches Agentic Video Understanding in Flash Models
Developers gain the ability to dynamically search, scan, and inspect long videos in Gemini Flash models.
Rohan Doshi, Gemini multimodal PM at DeepMind, posted the launch of agentic video understanding in Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The official Google account confirmed the update and shared a 58-second demo video plus a scatter plot on the 1H-VideoQA benchmark. The feature lets models scrub, zoom, edit, and resample frames instead of ingesting video statically. Posts note an attached chart comparing tokens per query and accuracy, with 3.7 Flash ranking at the frontier. LangChain replied that agents can now use the capability.
Combined views
524.8K
6 posts, first seen 13h ago
