Gemini API Introduces Agentic Video Processing
Google engineers announce tool-based video processing in Gemini API that cuts tokens by up to 88%.
Google DeepMind engineer Omar Sanseviero and Google engineer Logan Kilpatrick announced Agentic Video in the Gemini API. The feature lets the model use tools to intelligently process long videos according to the prompt, including transcription analysis and FPS adaptation. This approach reduces token consumption by up to 88 percent compared to full processing, while lowering latency and improving accuracy. The capability is available with recent models such as 3.7 Flash and can be controlled on a per-video basis through the API.
Agentic video understanding is here🔥 Rather than spending 200k tokens processing a long video, Gemini can now use tools to intelligently process them based on the prompt (analyze transcription, adapt FPS, etc.) leading to 88% fewer tokens, lower latency, and higher accuracy!
Gemini API Introduces Agentic Video Processing
Google engineers announce tool-based video processing in Gemini API that cuts tokens by up to 88%.
Google DeepMind engineer Omar Sanseviero and Google engineer Logan Kilpatrick announced Agentic Video in the Gemini API. The feature lets the model use tools to intelligently process long videos according to the prompt, including transcription analysis and FPS adaptation. This approach reduces token consumption by up to 88 percent compared to full processing, while lowering latency and improving accuracy. The capability is available with recent models such as 3.7 Flash and can be controlled on a per-video basis through the API.
