Gemini API Introduces Agentic Video Processing
Google engineers announce tool-based video processing in Gemini API that cuts tokens by up to 88%.
TLDR
Google DeepMind engineer Omar Sanseviero and Google engineer Logan Kilpatrick announced Agentic Video in the Gemini API. The feature lets the model use tools to intelligently process long videos according to the prompt, including transcription analysis and FPS adaptation. This approach reduces token consumption by up to 88 percent compared to full processing, while lowering latency and improving accuracy. The capability is available with recent models such as 3.7 Flash and can be controlled on a per-video basis through the API.
Combined views
140.4K
2 Sources, first seen 29d ago
