Introducing agentic video understanding with Gemini
Sep 01, 2026
Our new agentic feature for video analysis cuts token consumption by up to 88%, reduces costs by up to 66%, and boosts quality by up to 7%.
Rohan Doshi
Senior Product Manager, Google DeepMind
Mario Lučić
Research Director, Google DeepMind
Listen to article
[[duration]] minutes
This content is generated by Google AI. Generative AI is experimental
Today, we’re launching agentic video understanding across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.
The feature is available today for video uploads and YouTube videos via the...
Copyright of this story solely belongs to blog.google. To see the full text click HERE