Introducing agentic video understanding with Gemini

https://storage.googleapis.com/gweb-uniblog-publish-prod/images/agentic-video___keyword__blog-header.width-1300.png

Sep 01, 2026

Our new agentic feature for video analysis cuts token consumption by up to 88%, reduces costs by up to 66%, and boosts quality by up to 7%.


Rohan Doshi

Senior Product Manager, Google DeepMind

Mario Lučić

Research Director, Google DeepMind


Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.

The feature is available today for video uploads and YouTube videos via the...

Copyright of this story solely belongs to blog.google. To see the full text click HERE